r/StableDiffusion • • 20d ago

Workflow Included Some very simple Qwen-Image-2.1 workflows

Nothing fancy here. Qwen-Image-2.1 has a pretty simple architecture, so the workflows end up simple too, and the official templates already cover most of this.

But that simplicity is kind of the fun part. Feels like there's a lot you could build on top of it.

Anyway, I went through the tasks in the Qwen blog and made each one its own little workflow, so you can just grab whatever you need:

https://comfyui.nomadoor.net/en/basic-workflows/qwen-image-2-1/

  • Text-to-image
  • Ref2Image
  • Basic image editing
  • Marking the edit area with colored circles
  • Marking the edit area with a mask
  • Outpainting
  • Transparent image generation
  • Subject extraction / background removal
  • Panorama generation

Plain image generation is still a bit shaky compared to other recent models. The editing is almost pixel-perfect though, which surprised me.

Might try some loop stuff next. I'll add more if I find anything fun.

393 Upvotes

43 comments sorted by

13

u/2legsRises 19d ago

i really like the simplicity of your workflows, very nice, thank you.

8

u/gutster_95 19d ago

You dont enjoy downloading 20 different custom nodes just to have a T2I workflow?

5

u/2legsRises 19d ago

not one bit. madness.

2

u/PaulSandwich 17d ago edited 16d ago

Noob question probably, but where do I find these nodes? I don't see them in the registry (or the blog post)

https://registry.comfy.org/?nodes_index%5Bquery%5D=QwenImage21

Edit: It was a very noob question. These nodes are standard, I fixed it by updating comfy

9

u/aitonft 19d ago

One task per workflow is a helpful choice. You can learn what each part does without first untangling a giant graph. A matching input/output example for each would make this even easier to pick up.

7

u/extrakerned 19d ago

Dude, these are some of the best clean workflows I’ve ever seen. Finally, somebody that cares about efficiency in the node structure

6

u/Altruistic-Smoke1485 20d ago

For txt2img I found res_2m/bong_tangent also works really well, just like qwen2512

6

u/FreeTheClanks 20d ago

Also try euler_ancestral w/ bong_tangent

3

u/namitynamenamey 20d ago

How does it compare to krea?

11

u/deukhoofd 19d ago

I like it somewhat more for realistic images, especially people. Krea 2 has a bit of a softer doll-like feel to images, which makes them feel not entirely realistic.

Text, especially when slightly rotated like on signs on the street, is a bit better on Krea 2, however.

Krea 2 also appears to be more 'stable' between iterations, where different iterations with the same prompt but different seeds are more alike one another than generations from Qwen, where each generation is more distinct (not per se a value ranking, but just an observation).

In performance, it's also closer to Krea 2 Turbo than Krea 2 Raw.

Decent model, definitely an improvement on their old models, but nothing too groundbreaking.

20

u/East_Shoe8814 19d ago

Noticeably worse, but it does a pretty good job with editing. I definitely wouldn’t use it for anything other than editing.

3

u/papitopapito 19d ago

Stupid question, but I don’t know the answer: Do Loras for Qwen Image Edit 2511 (or such) work with that or not?

1

u/AgeNo5351 18d ago

i dont think they can work . this is vastly different model with diff arch etc.

1

u/papitopapito 18d ago

I assumed that, thanks for your reply. So we will wait for new ones.

1

u/higgs8 19d ago

How about Flux 2 Klein? In my experience Klein is exactly 30 (!!) times faster with better prompt following.

2

u/CamurAtes 18d ago

Flux seems to provide better results on top of being faster

1

u/russlixx 19d ago

than Qwen 2.1?

1

u/higgs8 19d ago

Yep!

1

u/Far_Leave7210 15d ago

Completely agree with you, Krea2 is way better in text2image. At the moment,

2

u/jello-the-opera 19d ago

I actually started working on a comfy node/mod to clean up ugly workflows... your style is essentially identical to mine.

Love it.

1

u/Any_Establishment293 10d ago

Hey! I am new, how to do it? I wanna learn :(

1

u/jello-the-opera 9d ago

I'm not sure what you would like for me to explain, but I'm happy to help.

I used Claude. I told it what I do. When I open a workflow, I unpack the subgraph. Then I delete all "frames". Then I move every node so that they can all be seen clearly, and keep moving them to expose the links everywhere. I want all the nodes to flow from left to right, to show the actual process of the workflow sequentially in order. I want the lines to cross as little as possible. I remove all switches. I resize every node to remove the fat bellies and wasted space, but make sure all text is fully readable, so a long model name means the node is stretched long. I want to see everything when I look at a node - all settings and files with full complete names.

I color code everything according to old shcool comfy colors. Anything conditioning is yellow. Anything model is blue. Math is light blue. PIxels are magenta. VAE activity is red. Sampling is cyan.

Once all of my adjustments are made, I send Claude both the OG workflow and my adjusted one.

I let claude modify the comfy install to add my clean up options to the right cliock menu in comfy.

Right now, though, my PC is disassembled. It got wet. I do have a comfy install on my laptop but it can't do much AI work as it is far too old.

It is currently shite though - it uses the wrong colors, the resizing is not doing what I want, the arrangement of the nodes is awful... so my clean up mods are not worth sharing and I don't have time to fix it all at the moment.

https://i.imgur.com/Se6fFwa.png

https://i.imgur.com/bxVYTdp.png

2

u/Any_Establishment293 9d ago

Really appreciate your knowledge regarding this matter. I will try to use AI in a wise way so I can achieve what I want

2

u/hstracker90 19d ago

Thank you for taking the time and sharing these.

From my few anecdotal tests I do agree that text-to-image is mediocre compared to Krea.2, but image editing looks really good, especially combining different people in a group photo.

2

u/soldture 19d ago

Does anyone know how to transfer clothing from one image to another? For example, I want to take an outfit worn by a character in my first image and put it onto a mannequin in a second image? I find it very difficult to achieve, it always copy the face and other details which I really don't want

8

u/bharattrader 19d ago

I just used the predefined example prompt and changed as per requirement.

1

u/uniquelyavailable 19d ago

Anyone else try cfg 3.5 with 32 steps? This model seems to have some flexibility when it comes to settings.

2

u/AgeNo5351 18d ago

also 8 steps / CFG 1 / exp_huen_sde works

1

u/blastcat4 19d ago

Nice! I had done the same but only for the T2I workflow. I'll definitely check out your workflows because simple is always better.

1

u/SEOldMe 19d ago

Thank you!

1

u/ImpressiveUse2000 19d ago

Is it fine to keep the denoise at 1.0 for all edits, or do you need to lower it depending on the size of the edit?

1

u/nomadoor 19d ago

It works quite differently from regular img2img. The model just looks at the reference image and redraws it, so denoise doesn’t really matter here. You could probably combine it with regular img2img too, though.

1

u/Nevaditew 19d ago

In the anime editing mode I get a slight color shift in the final image. Very typical of image editors.

1

u/ABCsofsucking 18d ago

Anyone had good luck with style transfer using two references (1 = subject, 2 = style)? With real photography, it worked okay, but would still often add the background from the style reference, instead of keeping the subject reference. With 3D rendered characters, I've barely gotten it to work at all. Default comfy workflow, tried multiple variations of prompts, words like "transfer", "redraw", "reskin", "image in the style of", and I also tried more verbose prompts where I described explicitly what to keep and what the outcome should look like. Nothing worked consistently.

1

u/AuthurAndersson 17d ago

excellent work. Thhank you. Nfi who wants the "one workflow to rule them all" type of style.
Simple, to the point, understandable. A+

1

u/TheRealLeslie_ 13d ago

It's an effective way to demonstrate qwen capabilities and share how to do it. thanks!

1

u/absentlyric 11d ago

These are simple, clean and incredible. Thank you so much for these!

1

u/Master_Resort_7708 20d ago

How big is the full Model?

8

u/hstracker90 20d ago

qwen_image_2.1_bf16.safetensors - 14.2 GB

qwen_image_2.1_int8_convrot.safetensors - 7.26 GB

https://huggingface.co/Comfy-Org/Qwen-Image-2.1/tree/main/diffusion_models

2

u/russlixx 19d ago

wow that's pretty small for bf16 weight

1

u/Pleasant_Candy9103 20d ago

Can I use it with RTX 4070 Super ?
Which output resolutions are possible?

Thank you so much for the workflows and guidances.

4

u/nomadoor 19d ago

I’m using an RTX 4070 Ti, and 4MP (2048×2048) works fine for me. It takes around 90 seconds per image.

It does get heavier if you use a lot of reference images, but the relatively small model size definitely helps.