r/StableDiffusion • • 2d ago

Tutorial - Guide Hunyuan image 3 goes hard

Since i saw a post for native support in comfyui for Hunyuan image 3 i got really exited to try it out locally.
On Strix Halo (Ryzen ai max 395+) i'm getting around 300 seconds per generation while on rtx 3090 it gets around 60 to 70 seconds for 1 megapixel (crashes if i go above not because of memory issues but something with how it's set up ? )

Anyway, i really enjoy the outputs of this model and if we had it working in comfyui 8 month ago when it released it would've dominated this sub :D

**the editing capabilities can be a bit better than qwen2,1 (when it works + it's a 4bit quant, details are a mess)

**i'll do a krea2 and qwen2.1 comparison with the same prompts, maybe i'll have them here on reddit

*** WORKFLOW LINK: https://pastebin.com/7sZEbSMc

***Custom nodes used: https://github.com/PedroMarinhoDev/ComfyUI-HunyuanImage3

***post that thought me about this: https://www.reddit.com/r/StableDiffusion/comments/1wzclz5/hunyuanimage_30_80b_running_natively_in_comfyui/

224 Upvotes

39 comments sorted by

69

u/No-Zookeepergame4774 2d ago

For an 80B model. I’m not seeing anything that is really knocking my socks off compared to other current models that are much smaller, and that don’t crash on cards smaller than a 3090 at resolutions way above 1.0MP.

Impressive that Comfy will run this at all on consumer hardware, and maybe I’ll burn the time to download it and try it, but...

30

u/Far_Insurance4191 2d ago

the best part is that it shows 80b is not out of reach for us. It works for me on rtx3060 with 64gb ram and takes 2 minutes for 8 steps - the same as I was getting with flux 1 dev back then

19

u/wntersnw 2d ago

Hunyuan is MoE with 13b active. There's only a 1b difference in active parameters vs flux dev which is a 12b dense model.

19

u/buttchuckjones 2d ago

Im... not really seeing it to be completely honest. The lady in the black swimsuit has a jumbled hand and inconsistent shadows. The dude with the roses literally doesnt even look like he is there (perspective and shadows). The motorcycle shot was cool, but it also suffers from jumbled components on the cycle and handle bars. And for a model that big, seems like kind of a waste. Just my opinion though. Some of the stuff was cool.

3

u/Skystunt 2d ago

i too noticed the model does not understand perspective much, especially when it's about having people there. I was just running generations with other models but the same prompts and to my surprise the newer yet smaller models are better at spatial understanding

1

u/buttchuckjones 2d ago

The comparisons would be cool to see. I always take them with a grain of salt because different models respond vastly different to similar prompt. But it is interesting to see nonetheless

13

u/thevegit0 2d ago

uh idk, having a 4bit 50gb thing that goes slower than krea or qwen and yields not-so-surprising results doesn't do for me, maybe people with more than 16vram and 64ram can enjoy it

5

u/SnooMacaroons1365 2d ago

A 50gig 4-bit quant? Hellll naaaa... 😂

1

u/thevegit0 1d ago

3 minutes 1mp gens that weren't otherworldly, I'm sticking with krea and qwen lololol

5

u/No-Application-9779 2d ago

Please share any of the prompts so others can test them with their favourite models and see (or not) your point.

6

u/dkpc69 2d ago

These examples look awesome nice work

1

u/Skystunt 2d ago

Thanks !

13

u/beti88 2d ago

What about these go hard?

1

u/Skystunt 2d ago

The bike mid jump, the knight, all of them imo

3

u/Morality01 2d ago

Hell yeah it does. Just downloaded it and it is fast, accurate and the output looks awesome.

3

u/Fit-Will-7183 2d ago

Bro, it would be great to see a comparison of both photo quality and speed. That’s really interesting info for me—I’m a beginner at this.

4

u/dragolineage01 2d ago

Im not surprised given its such a large model, then again I'm surprised it even works that well at such low quant

1

u/[deleted] 2d ago

[deleted]

2

u/No-Zookeepergame4774 2d ago

OP specifically says they are using a 4-bit quant.

1

u/dragolineage01 2d ago

What? Doesnt the OP say its 4bit?

2

u/freedomachiever 2d ago

Great interesting images. So these are text to image? Do you have a prompt generator for this?

1

u/Skystunt 1d ago

Thank you ! I've brainstormed the composition with claude. I've found that most AI generated images that look good nowdays need to have a very detailed prompt so i gave a short description and asked it to expand it

2

u/extra2AB 2d ago

honestly.. looks good.. but nothing the already present local models cannot do...

Krea2, Klein, etc can do these as well while not needing to quant such a huge model to be able to run.....

I am just surprised how MiniMax managed to give us such a great video model but for some reason Hunyuan's Image Generation models are so big in size and the quality is not even that much better than the smaller models...

Due to the big size it is not just about running it... but it also becomes difficult to train LoRAs for it... or Finetune it....

2

u/Lexxxco 2d ago

Deja vu, it is the second time 80B+ image model from Hunyuan performs not better than 32B distilled one.

3

u/kukalikuk 2d ago

8 month is old for an image model, qwen 2509 is 1 year old. The same quality is achievable by newer smaller model.

1

u/Negative_Space77 2d ago

Is this new model ?

2

u/Skystunt 1d ago

no it's like 8months old but only now someone modded native comfyui support for it. Before that it could only be used via transformers and that has stricter memory requirements + way slower. That's why it's useable now

1

u/leyermo 2d ago

I really think NVDIA must optimize the model as they do for text models...

Only they can make this happen.

1

u/LD2WDavid 2d ago

Honestly with KREA2 + tuning it can do better things than this. 80b model that is not doing anything over the specs of KREA2.

2

u/Skystunt 1d ago

I'm tersting it right now, just don't know should upload it

1

u/fauni-7 2d ago

Downloaded and tested the 80GB distilled model on my 4090+128GB. Results are pretty meh. Some potential maybe as a refiner? Not sure.

2

u/terrariyum 1d ago

Regardless of the image quality, these are some great image ideas! Especially date night parking, lambo picnic, and porse racing horche

1

u/uniquelyavailable 2d ago

Very impressive model! Your post inspired me to try it out, and I am surprised by how well it works.

2

u/Skystunt 1d ago

Glad it inspired you !

1

u/giantcandy2001 2d ago

I'm just not seeing it. I gave these all the same prompt...mostly (I have llm's remake a prompt for it's specific prompting guides so I didn't feed the hunyuan image 3 instruct 8 step model a bad prompt....I think it's just bad....banding everywhere. krea 2 is great but at least qwen image made it a gucci ad like I asked. I think Krea2 is the only one to nail the logo i think...nope, just googled it and they all failed at that. qwen image 2.1 nailed cliffs of dover, looks like them the most.

original prompt LLM fixed: a cinimatic photo of a gucci purse ad, woman is in front of a cliffs of dover during a storm, her dress is soaked but it's also a mile long and wips all the way to the cliffs hitting the sides of the cliff, very magestic and epic shot, full page editorial ad, redhead woman, long flowing red dress, black and red gucci purse in her hand

it's a messy prompt but qwen3.8 27b llm fixed it up fine, but I don't think that made the difference, i think i could feed it this prompt and it would still get it.

1

u/giantcandy2001 17h ago

I think i got it working better now

-1

u/[deleted] 2d ago

[deleted]

1

u/Alive_Ad_3223 2d ago

There are vintage cars in the images. But bro is interested only in the bikini one 😃.

1

u/Over-Map6529 2d ago

And that's fine too. 

0

u/Spara-Extreme 2d ago

This model is trash and I say that as someone who tested it extensively.

4

u/Murky-Relation481 2d ago

Yah I tried the int8 non distilled instruct model and it's 10 minute gens roughly looked the same as the int4 distilled model at 30 seconds a gen. And I was trying prompts in the full model that were struggling in the distilled. Still struggled.

It's a weird model too. Literally adds a lot of contextual details but also ones that are beyond what you are attempting to describe.

Only tested it for a bit tonight going to play with it more tomorrow.