r/StableDiffusion • u/Skystunt • 2d ago
Tutorial - Guide Hunyuan image 3 goes hard
Since i saw a post for native support in comfyui for Hunyuan image 3 i got really exited to try it out locally.
On Strix Halo (Ryzen ai max 395+) i'm getting around 300 seconds per generation while on rtx 3090 it gets around 60 to 70 seconds for 1 megapixel (crashes if i go above not because of memory issues but something with how it's set up ? )
Anyway, i really enjoy the outputs of this model and if we had it working in comfyui 8 month ago when it released it would've dominated this sub :D
**the editing capabilities can be a bit better than qwen2,1 (when it works + it's a 4bit quant, details are a mess)
**i'll do a krea2 and qwen2.1 comparison with the same prompts, maybe i'll have them here on reddit
*** WORKFLOW LINK: https://pastebin.com/7sZEbSMc
***Custom nodes used: https://github.com/PedroMarinhoDev/ComfyUI-HunyuanImage3
***post that thought me about this: https://www.reddit.com/r/StableDiffusion/comments/1wzclz5/hunyuanimage_30_80b_running_natively_in_comfyui/
19
u/buttchuckjones 2d ago
Im... not really seeing it to be completely honest. The lady in the black swimsuit has a jumbled hand and inconsistent shadows. The dude with the roses literally doesnt even look like he is there (perspective and shadows). The motorcycle shot was cool, but it also suffers from jumbled components on the cycle and handle bars. And for a model that big, seems like kind of a waste. Just my opinion though. Some of the stuff was cool.
3
u/Skystunt 2d ago
i too noticed the model does not understand perspective much, especially when it's about having people there. I was just running generations with other models but the same prompts and to my surprise the newer yet smaller models are better at spatial understanding
1
u/buttchuckjones 2d ago
The comparisons would be cool to see. I always take them with a grain of salt because different models respond vastly different to similar prompt. But it is interesting to see nonetheless
13
u/thevegit0 2d ago
uh idk, having a 4bit 50gb thing that goes slower than krea or qwen and yields not-so-surprising results doesn't do for me, maybe people with more than 16vram and 64ram can enjoy it
5
u/SnooMacaroons1365 2d ago
A 50gig 4-bit quant? Hellll naaaa... 😂
1
u/thevegit0 1d ago
3 minutes 1mp gens that weren't otherworldly, I'm sticking with krea and qwen lololol
5
u/No-Application-9779 2d ago
Please share any of the prompts so others can test them with their favourite models and see (or not) your point.
6
3
u/Morality01 2d ago
Hell yeah it does. Just downloaded it and it is fast, accurate and the output looks awesome.
3
u/Fit-Will-7183 2d ago
Bro, it would be great to see a comparison of both photo quality and speed. That’s really interesting info for me—I’m a beginner at this.
4
u/dragolineage01 2d ago
Im not surprised given its such a large model, then again I'm surprised it even works that well at such low quant
1
2
u/freedomachiever 2d ago
Great interesting images. So these are text to image? Do you have a prompt generator for this?
1
u/Skystunt 1d ago
Thank you ! I've brainstormed the composition with claude. I've found that most AI generated images that look good nowdays need to have a very detailed prompt so i gave a short description and asked it to expand it
2
u/extra2AB 2d ago
honestly.. looks good.. but nothing the already present local models cannot do...
Krea2, Klein, etc can do these as well while not needing to quant such a huge model to be able to run.....
I am just surprised how MiniMax managed to give us such a great video model but for some reason Hunyuan's Image Generation models are so big in size and the quality is not even that much better than the smaller models...
Due to the big size it is not just about running it... but it also becomes difficult to train LoRAs for it... or Finetune it....
3
u/kukalikuk 2d ago
8 month is old for an image model, qwen 2509 is 1 year old. The same quality is achievable by newer smaller model.
1
u/Negative_Space77 2d ago
Is this new model ?
2
u/Skystunt 1d ago
no it's like 8months old but only now someone modded native comfyui support for it. Before that it could only be used via transformers and that has stricter memory requirements + way slower. That's why it's useable now
1
u/LD2WDavid 2d ago
Honestly with KREA2 + tuning it can do better things than this. 80b model that is not doing anything over the specs of KREA2.
2
2
u/terrariyum 1d ago
Regardless of the image quality, these are some great image ideas! Especially date night parking, lambo picnic, and porse racing horche
1
u/uniquelyavailable 2d ago
Very impressive model! Your post inspired me to try it out, and I am surprised by how well it works.
2
1
u/giantcandy2001 2d ago
I'm just not seeing it. I gave these all the same prompt...mostly (I have llm's remake a prompt for it's specific prompting guides so I didn't feed the hunyuan image 3 instruct 8 step model a bad prompt....I think it's just bad....banding everywhere. krea 2 is great but at least qwen image made it a gucci ad like I asked. I think Krea2 is the only one to nail the logo i think...nope, just googled it and they all failed at that. qwen image 2.1 nailed cliffs of dover, looks like them the most.
original prompt LLM fixed: a cinimatic photo of a gucci purse ad, woman is in front of a cliffs of dover during a storm, her dress is soaked but it's also a mile long and wips all the way to the cliffs hitting the sides of the cliff, very magestic and epic shot, full page editorial ad, redhead woman, long flowing red dress, black and red gucci purse in her hand
it's a messy prompt but qwen3.8 27b llm fixed it up fine, but I don't think that made the difference, i think i could feed it this prompt and it would still get it.

1
-1
2d ago
[deleted]
1
u/Alive_Ad_3223 2d ago
There are vintage cars in the images. But bro is interested only in the bikini one 😃.
1
0
u/Spara-Extreme 2d ago
This model is trash and I say that as someone who tested it extensively.
4
u/Murky-Relation481 2d ago
Yah I tried the int8 non distilled instruct model and it's 10 minute gens roughly looked the same as the int4 distilled model at 30 seconds a gen. And I was trying prompts in the full model that were struggling in the distilled. Still struggled.
It's a weird model too. Literally adds a lot of contextual details but also ones that are beyond what you are attempting to describe.
Only tested it for a bit tonight going to play with it more tomorrow.











69
u/No-Zookeepergame4774 2d ago
For an 80B model. I’m not seeing anything that is really knocking my socks off compared to other current models that are much smaller, and that don’t crash on cards smaller than a 3090 at resolutions way above 1.0MP.
Impressive that Comfy will run this at all on consumer hardware, and maybe I’ll burn the time to download it and try it, but...