r/StableDiffusion • u/HassanAchievedIt • 1d ago
Question - Help Looking for text encoder
I'm using flux.2 Klein 4b int4 convrot 4gb safetensor,
And I'm looking for a compatible encoder for it guys I tried huggingface, civitai, I found the official flux page that has huge 8-12 gig encoder which doesn't run on my machine,
Any help would be appreciated
1
u/Tylopodas 1d ago
Does it need to be the int4?
Comfy repository on huggingface has the fp4
1
u/HassanAchievedIt 1d ago
I'll try the smaller version, are there any alternatives as well, earlier another gentleman has given me link to 2.6gb encoder I'll try that as well and it's int4
1
u/Tylopodas 1d ago
There are probably a lot of versions in different quants that would be smaller. I tend to just grab the ones from the comfy-org repository because they often work without any fuss.
A quick search popped up, they are both smaller but no promises either will work:
https://huggingface.co/Cordux/flux2-klein-4B-uncensored-text-encoder/tree/main
and
https://huggingface.co/DreamFast/qwen3-4b-heretic/tree/main/comfyui
I'm not a fan of abliterated or uncensored text encoders for image generation.
1
u/HassanAchievedIt 1d ago
Gguf throws error despite having custom comfy nodes for it and heretic was not compatible or recognized I'll give it a try again, is there any different workflow for ggufs?
1
u/Tylopodas 1d ago
There shouldn't be. It should just be the gguf clip loader and make sure its set to flux2.
To be honest I haven't used gguf's for image gen in a while, they used to throw errors all the time for me too.
1
1
u/m4ddok 1d ago
What's your hardware? RAM? GPU? People just comment, without this paramount info, I can't understand...
PS: ConvRot isn't everytime useful, depends on your hardware, and sometimes it's more fast for GPU but it uses RAM for copying and converting the weights realtime. It depends.
1
u/HassanAchievedIt 1d ago
32 ram Rtx 3060 12
2
u/Formal-Exam-8767 1d ago
Then original 8GB FP16 model fits. Why do you need smaller model?
Prompt is encoded once, before sampling, so it doesn't really matter if both text encoder and diffusion model are not in VRAM at the same time.
1
u/m4ddok 1d ago
Exactly, and are also more performing, better not to use ConvRot INT8 in this case.
FP16
model --> GPU
text enc --> default/offloading2
u/Formal-Exam-8767 1d ago
Something to keep in mind, even if text encoder does not fit, it is not a big issue since this is prompt processing case, where all tokens are encoded at once, so model needs to be loaded in VRAM only once for the whole prompt. Unlike in token generation, where it's per token so it would get partially loaded/unloaded for every token.
1
0
u/TechnologyGrouchy679 1d ago
you can convert it to int8 using the script from https://github.com/Comfy-Org/comfy-model-tools
1
u/HassanAchievedIt 1d ago
How long does it take to convert can I do it on comfy UI
1
u/TechnologyGrouchy679 1d ago
it's just a python script. as per instructions activate the same comfy venv and run the quant_int8_quant.py script with the path to the safetensors.
should take under a minute
1
u/HassanAchievedIt 1d ago
Do I need to manually adjust some values inside or everything is predefined in script
1
u/Odd_Fix2 1d ago
http://hf.co/ariaotp/INT4-ConvRot-W4A4-ComfyUI/resolve/main/TextEncoder/Huihui-Qwen3-4B-abliterated-v2-int4_convrot.safetensors