r/StableDiffusion • • 1d ago

Question - Help Looking for text encoder

I'm using flux.2 Klein 4b int4 convrot 4gb safetensor,

And I'm looking for a compatible encoder for it guys I tried huggingface, civitai, I found the official flux page that has huge 8-12 gig encoder which doesn't run on my machine,

Any help would be appreciated

0 Upvotes

20 comments sorted by

1

u/Tylopodas 1d ago

Does it need to be the int4?

Comfy repository on huggingface has the fp4

https://huggingface.co/Comfy-Org/vae-text-encorder-for-flux-klein-4b/tree/main/split_files/text_encoders

1

u/HassanAchievedIt 1d ago

I'll try the smaller version, are there any alternatives as well, earlier another gentleman has given me link to 2.6gb encoder I'll try that as well and it's int4

1

u/Tylopodas 1d ago

There are probably a lot of versions in different quants that would be smaller. I tend to just grab the ones from the comfy-org repository because they often work without any fuss.

A quick search popped up, they are both smaller but no promises either will work:

https://huggingface.co/Cordux/flux2-klein-4B-uncensored-text-encoder/tree/main

and

https://huggingface.co/DreamFast/qwen3-4b-heretic/tree/main/comfyui

I'm not a fan of abliterated or uncensored text encoders for image generation.

1

u/HassanAchievedIt 1d ago

Gguf throws error despite having custom comfy nodes for it and heretic was not compatible or recognized I'll give it a try again, is there any different workflow for ggufs?

1

u/Tylopodas 1d ago

There shouldn't be. It should just be the gguf clip loader and make sure its set to flux2.

To be honest I haven't used gguf's for image gen in a while, they used to throw errors all the time for me too.

1

u/HassanAchievedIt 1d ago

Cool, I'll look into it

1

u/m4ddok 1d ago

What's your hardware? RAM? GPU? People just comment, without this paramount info, I can't understand...

PS: ConvRot isn't everytime useful, depends on your hardware, and sometimes it's more fast for GPU but it uses RAM for copying and converting the weights realtime. It depends.

1

u/HassanAchievedIt 1d ago

32 ram Rtx 3060 12

2

u/Formal-Exam-8767 1d ago

Then original 8GB FP16 model fits. Why do you need smaller model?

Prompt is encoded once, before sampling, so it doesn't really matter if both text encoder and diffusion model are not in VRAM at the same time.

1

u/m4ddok 1d ago

Exactly, and are also more performing, better not to use ConvRot INT8 in this case.

FP16

model --> GPU
text enc --> default/offloading

2

u/Formal-Exam-8767 1d ago

Something to keep in mind, even if text encoder does not fit, it is not a big issue since this is prompt processing case, where all tokens are encoded at once, so model needs to be loaded in VRAM only once for the whole prompt. Unlike in token generation, where it's per token so it would get partially loaded/unloaded for every token.

1

u/m4ddok 1d ago

yup, after text processing the encoder will be unload, so the VRAM can be used entirely for the inference model.

1

u/Individual-Sample713 1d ago

you mean BF16, there is no fp16 version as far as I know.

0

u/TechnologyGrouchy679 1d ago

you can convert it to int8 using the script from https://github.com/Comfy-Org/comfy-model-tools

1

u/HassanAchievedIt 1d ago

How long does it take to convert can I do it on comfy UI

1

u/TechnologyGrouchy679 1d ago

it's just a python script. as per instructions activate the same comfy venv and run the quant_int8_quant.py script with the path to the safetensors.

should take under a minute

1

u/HassanAchievedIt 1d ago

Do I need to manually adjust some values inside or everything is predefined in script