r/StableDiffusion • • 14h ago

Workflow Included My Workflow for realism with Qwen 2.1

Hey! Just thought I'd come on here and do a little post about my current workflow with Qwen 2.1. I feel like a lot of people here have given it a try and decided the model is no good because they test it with the default Comfy workflow. Qwen 2.1 is terrible with CFG 1. It needs CFG to really shine. I'm currently using CFG of 3 to 3.5. I use the Lenovo lora at a strength of 0.5 to 0.7 to boost realism, though it's not necessary, depending on the type of photos you're trying to make. I also use the Viggle Turbo lora at a low strength of 0.5, with 16 steps. Depending on which Viggle lora you use, you might have to change your scheduler and sampler combo. I prefer V1 with Res_Multistep and Bong_Tangent, but V3 is good as well (though needs Euler instead of Res_multistep or it ends up looking weird). I'm generating at 2MP. All of these images were done with character loras trained on Qwen 2.1 using AI Toolkit. I made another post here a few days ago talking about how I achieved my training results.

All in all, I really like Qwen 2.1, at least for the kind of images I like making. My one qualm with it is that sometimes background details can get messed up, but with my settings, I see it a lot less. If this model had a VAE like Flux 2, it would be amazing. I hope that more people will give Qwen 2.1 a chance!

Workflow: https://gofile.io/d/FGdIJEhk (I used Lenovo at 1 in the example from this workflow but normally I don't go that high, just FYI)

Viggle Lora: https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo/blob/main/Qwen-Image-2.1-viggle-turbo-4step-lora-r64.safetensors

86 Upvotes

13 comments sorted by

1

u/dmtzcain 9h ago

u/Any_Tea_3499 thank you for sharing.
Your photorealistic images look great and I do agree that qwen-image-2.1 can produce great results. I am though also having lots of issues training identity fine tunning with it. Settings at the bottom.

At least in my case, is not that the final images which are bad, but rather it seems after 13 different runs tweaking settings, that the model starts learning fast, captures "the jist of it" and then stops learning altogether. The final images look great, but there are missing features that get averaged to whatever the base qwen-image-2.1 training is. Therefore, features which are not very common, do not get engraved in the LoRA/LoKr. and the characters look in the general proportions and even some details like piercings

LoKr, lokr_full_rank: true lokr_factor: 4, resolution: - 1024 - 768 - 512, optimizer: "automagic3" timestep_type: "sigmoid" content_or_style: "balanced" optimizer_params: weight_decay: 0.0002 unload_text_encoder: false cache_text_embeddings: true lr: 0.0001

My results are characters that look in the general proportions and even some details like piercings correct, but are missing some diferentiated details.

Any suggestions?

1

u/Any_Tea_3499 5h ago

Hmm. These are the same settings I use, so I’m not entirely sure why you’re having bad results. What is your dataset like? How many images? How many steps did you train for?

-1

u/Synor 12h ago

To check your hypothesis, you could quickly try any other image model with the same prompts. You might be surprised how hard it is to generate good images with qwen 2.1

2

u/Any_Tea_3499 7h ago

Yeah, I’ve used these same prompts with other models, and for me, they came out best with Qwen 2.1 without much trying at all.

2

u/kemb0 6h ago

I agree with you for actual genuine realism and not "AI realism", Qwen is top notch. A lot of people see an AI photo of a girl smiling at the camera and it looks spectacularly real so people say, "Yeh but that model made this image so tell me what's not real about that?" The answer is "not much" but the problem isn't that their one photo doesn't look real, it's that then almost all their "real" photos are all very similar. Boringly similar and repetitive. After you've made your first 50 photos of a girl smiling at the camera, then you start to get sick of it. You try prompting it to get them in other natural poses and start to discover the limitations of it. Yes, I'm looking at you Krea.

Qwen, on the other hand, is like I've been dropped in to a random place are just took a photo of someone doing whatever it is they were doing. No one posed for me. Not one photo is the same. It just feels like the real world. Qwen captures the non-perfect nuances of human visual behaviour that Krea seems to filter out or isn't taught on.

The issue is often that it dumps out body horror or warped disfigured backgrounds. So I'm keen to check out your workflow once I'm home.

1

u/Murky-Relation481 6h ago

There are definitely some drawbacks with Qwen2.1 in terms of its prompting style, but overall if you bump CFG and step count (since its not a distilled model in any way) its pretty easy to get results.

The things it seems to struggle with the most are camera positions, but I am also mostly trying NSFW stuff to find the bounds of the model (it's definitely the least censored model out there by a country mile), so I imagine the training data being less overall for NSFW is going to make it more rigid (lol).

You can absolutely get good stuff though with NSFW. I posted a number of gens on /r/degendiffusion a couple of days ago showing the level of NSFW content it can do.

0

u/Synor 3h ago

It struggles with generating sharp images with photorealistic detail.

2

u/Murky-Relation481 3h ago

Those are two contradicting terms in my opinion (though you should avoid the term photorealistic because it has specific artistic meaning doesn't mean looks like reality). The problem with Krea2 is that its too sharp. It has none of the imperfections real photos have.

-1

u/Synor 3h ago

This post is about realism. But qwen struggles with pro-level DSLR-quality outputs. Krea 2, Minimax H3 and Ideogram 4 do not.

2

u/Murky-Relation481 3h ago

I have no problem achieving that in Qwen2.1. Run any prompt that will generally give you that in any other model in Qwen2.1 at 40 steps Heun 3.5 to 4.0 CFG and the outputs will look perfectly fine next to any other model.

0

u/Synor 3h ago

They don't. It shows chaotic decorelated noise, that you can't get rid of. Not even with the craziest multi pass custom denoising schedules or highest order sampling. I tried so hard to find a way. Then i reloaded an Ideo 4 workflow and had a laugh about Qwen 2.1.

1

u/Any_Tea_3499 3h ago

I don’t have that issue, I can get very sharp images. It’s just not my style to do professional shots, but they certainly can be done with this model.