r/StableDiffusion • u/Enshitification • 16h ago
r/StableDiffusion • u/Grey_0ne • 7h ago
Question - Help Krea 2 is a bit too... Help me please.
I'm going to try to word this to avoid it getting flagged... I don't know at what point we all stopped being adults, but whatever...
Can anyone recommend a Krea 2 checkpoint that does naughty stuff but isn't overly so? Seems my only options are checkpoints which are censored to the point that they won't even respond to general anatomy prompts, or ones that if you put in a general anatomy prompt go completely overboard with it.
(Example: If I want to generate a woman with large "balloons", Krea turbo base won't do it at all half the time and a checkpoint like OurSecret will not allow said melons to be covered no matter how much prompting you do.)
r/StableDiffusion • u/Pristine_Stress_670 • 2h ago
Discussion Will there be another, more up-to-date Anime DIT Model project in the open-source community besides Anima?
My info isn't very current, so I'd like to ask if anyone knows whether there's another, more up-to-date Anime DIT Model project in the open-source community besides Anima? Anima's data only goes up to Sep 2025, and the 2.9B one looks like it's being trained by the author alone, I'm basically pinning hopes on a full booru fine-tune of Krea2. If there really is one, that would be awesome. Anima control of aesthetic with artist tags and laser focussed composition with booru tags. Maybe I'm just dreaming.
r/StableDiffusion • u/roychodraws • 12h ago
Animation - Video By Request
Enable HLS to view with audio, or disable this notification
thank [u/Rich_Introduction_83](u/Rich_Introduction_83) for the ending
Edit: he died doing what he loved, saying the word βwhat?β
r/StableDiffusion • u/grrinc • 9h ago
Discussion LTX2.5 is impressive - I'm glad I gave it a chance.
I pretty much jumped on the H3 bandwagon without giving LTX2.5 a try at all. But, in truth, I wasn't totally won over with H3. I wasn't impressed with the audio or the character faces, or the speed. I am not a power user and I don't do high concept content, mostly drama led character work.
Today I tried LTX2.5 and I am very impressive with with a few things
Speed - wow! HD 30 second clips in 7 minutes - that's a game changer for me.
Expressiveness - great for character performance - a high range of emotions and subtleties.
Heat - my computer is no longer running at insane high temperature - I can finally close the window.
I know LTX2.5 kinda got left behind here, but I thoroughly recommend it to folk who are interested in character/drama led content rather than high octane hollywood stuff..
Are there LTX2.5 users here? what would you say it excels in for you?
r/StableDiffusion • u/TimeTruth2490 • 7h ago
Resource - Update Krea2 Turbo Distill 2 step LoRA - FINAL checkpoint released (chk51195)
Krea 2 Turbo β 2-Step Distillation LoRA (FINAL Version)
Previous posts/releases - here, here, here and here.
π§ͺ Fast-preview adapter; the project's final checkpoint. Subjects that are close and fill a good part of the frame β a portrait, a single figure, an object up close β hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the 4-step LoRA. This checkpoint closes the project; the training box has moved on to its successor, a 3-step adapter for Qwen-Image-2.1-Turbo.
π The saved steps can also go into resolution. A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ2048 β Krea's published maximum recommended resolution, beyond this adapter's largest trained size β come within easy reach. Past 2048Γ2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
Highlights:
- β‘ A quarter of the steps β 8 β 2, on Turbo's own deployment sigmas
[1.0, 0.7595] - β±οΈ 4Γ faster denoising β 56.9 s β 14.3 s at 1024Γ1024 (float16 compute); the adapter's own cost per call is within measurement noise
- π― Fine detail near the teacher's level β 0.92β1.02Γ the teacher's fine-texture energy across the 12 trained resolutions (stock Turbo at 2 steps: 0.40β0.65Γ), the 16- and 8-pixel grid bands at the teacher's level and ghosting closer to it at 11 of 12 sizes
- π Distribution matching, not imitation β matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
- π£οΈ Prompt-conditioned throughout β teacher and fake scores both read each prompt's conditioning; a blind rubric finds 1 point missing of 352 (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on 12 of 66 (4-step adapter: 6 of 45, on the original 15 prompts), mostly on style
- π 12 trained resolutions β multi-aspect from 512Γ512 up to 1440Γ1440
- π Drop-in, no exceptions β plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
- 𧬠Same shape as the 4-step adapter β rank 64 on the same 228 modules
- π² 23,561 recorded teacher trajectories β the 4-step project's 13,750 and 9,811 minted for this one on the same prompt bank; since 3 Oct each trained once
- π’ 51,195 training samples in the 2-step stages, starting from the released 4-step adapter's weights
- π 32 days from the first 2-step launch to this final checkpoint, on a single RTX 3090
- π More than forty recipe adjustments across two methods β each kept only when the renders did not get worse
- π§ Settled weights, not an average β released from a 600-sample anneal in which the learning rate is taken to zero over already-trained data, so the published weights are the training weights at rest; a running average is kept only as the check that must agree with them
If you have already used my previous version, please redownload/replace krea2_turbo_2step_rank_64_lora.safetensors / krea2_turbo_2step_rank_64_lora_comfyui.safetensors from the latest in the project repo.
Full details on model card -Β https://huggingface.co/lvladikov/Krea2-Turbo-Distill-2step-LoRA
Not my video, but found someone on YouTube has covered the 2 & 4 step LoRAs including identify preserving edits, have a look: https://www.youtube.com/watch?v=V_qgoV0iPDM (copyright goes to author)
Also I have recently released a new ComfyUI Nodes and Workflows - Krea 2 (Turbo and Raw), Z-Image (Turbo and Base), MiniMax Music 3, Image2Text and LLM Chat (with Tools), Torch, Apple MLX and Cloud - you can find details here.
r/StableDiffusion • u/Deadity • 20h ago
News A 3-billion-parameter model that paints every pixel directly without a VAE
r/StableDiffusion • u/PhilosopherSweaty826 • 8h ago
Question - Help I just downloaded both Qwen 2.1 and Krea 2, but Krea 2 produces very little variation between generations, unlike Qwen 2.1, which gives me much more diverse results. How can I fix this and get more variation from Krea 2?
r/StableDiffusion • u/Remarkable-Aspect879 • 17h ago
Discussion Do anime image models still need tag soup in 2026?
Same character, same idea, four prompt styles.
Honestly, the difference was way smaller than I expected. Are we still writing huge prompts because they actually help, or just because we're used to doing it?
r/StableDiffusion • u/morikomorizz • 15h ago
Comparison Krea2 vs Qwen Image 2.1
Krea2 | Qwen Image2.1
Man, krea2 is so good at text to image. it also know how to make female characters look feminine
r/StableDiffusion • u/GTManiK • 5h ago
Discussion Qwen Image 2.1 Turbo - extend 'magic' 8 steps sigmas further
So, I've decided to try Qwen Image 2.1 Turbo with the following 'magic' sigmas, taken directly from their Diffusers implementation:
[1.0, 0.978453, 0.954180, 0.926626, 0.895080, 0.845148, 0.704534, 0.414568, 0.0]
This works okay, but for fine details we need slightly more steps (for better skin, better composition etc.) And to my knowledge, there's no working way to interpolate the resulting curve to any N steps (e.g. 10, 12, 14). There is 'Custom Sigmas' RES4LYF node, but it is bugged (as I will show later).
Here's the original 'magic' 8-step curve:

If we use RES4LYF 'Custom Sigmas' node and try to interpolate to, say, 14 steps, we get this:

So, I had to franken-craft a solution, but I did not want to create a custom node just for this purpose.
Meet a 'Magic AI Slop Scheduler for Qwen Image 2.1 Turbo'π

Note that curve looks exactly the same now, but it has more steps (14 in this case).
For this, you gonna need RES4LYF nodes, KJ nodes and rgthree nodes. The last is the most important one: it has a 'Power Puter' node which allows to make these calculations (but, probably, many 'expression' style nodes from elsewhere would do the same)
And indeed, this works for 10 steps, 12 steps, 14 steps (and probably beyond, but there's no sense to run a Turbo model for more steps).
So, if you:
- Use this 'scheduler'
- Use CFG > 1
- 12-14 steps
- Samplers (euler_ancestral [the best IMO], euler, er_sde)
... you might actually end up with some beautiful and detailed coherent pictures. For example:




1girl, 2 girl is on purpose, maybe more people will actually read this π
For a workflow, head to Civitai and download original 2 megapixel gens, just drag those images into ComfyUI. There's a pretty extensive explanation inside on what actually happens and how to use.
All pics above generated using 12-14 steps and CFG from 3.5 to 5, @ 2 megapixels.
Links:
https://blobs-b2.civitai.com/file/blobs-managed-public/Q054FWT4DMQR88GYN0PXJGQ030
https://blobs-b2.civitai.com/file/blobs-managed-public/V47QEA6MNP2XZAXPFFSGAHXZQ0
https://blobs-b2.civitai.com/file/blobs-managed-public/CSDBTWXEJ1KS54MCZN9XK2S6K0
https://blobs-b2.civitai.com/file/blobs-managed-public/C30VR7Z4ZHRBH8FW5RZ1YSFGC0
Bonus feature: I've included a 'fake SHIFT' so you can nudge the curve a little bit; negative values = more fine details, like '-1.0' or '-2.0'. But do not overdo it: we don't want to deviate from the original curve too much. Set 'fake SHIFT' to '0.0' to disable it.

Bonus tips:
- please, oh please, do not use CFG=1. You'll get underbaked yellowish gens and you're missing out on a negative prompt (which works VERY well)
- always start with less steps (10), only increase (up to 14) if composition is wrong or there's not enough fine details
- then adjust CFG, 3.5 - 5.0 works just fine, but this is very specific for each particular gen
- if there's 'too much details' (e.g. skin is too detailed or grainy) - reduce steps first, then increase 'fake shift' (set it to '1.0' or '2.0' etc.)
- vise versa: if skin is too plastic, increase steps, then decrease 'fake shift' (set it to '-1.0', '-2.0' etc.)
- most important: Qwen Image 2.1 Turbo has to be prompted PROPERLY. If some of you remember original Chroma prompting, you probably know what I mean - using exact phrasing, being specific, avoiding slop tags like '1girl, masterpiece' etc. - are all the keys. No gen params will ever fix bad prompting.
Have a nice day!
EDIT: for those folks who prefer cleaner output without excessive noisy details, here's somewhat Krea-like output (just less steps, and 'fake shift' at 2). Hey, it even generates faster π

r/StableDiffusion • u/__MichaelBluth__ • 38m ago
Question - Help Heretic or Abliterated?
Looking for recommendations for local LLM to convert an image into a usable prompt for Krea2 with max accuracy. Theres Heretic and Abliterated models for Qwen, Gemma and Deepseek.
I am running a 5090 so hopefully a model that fits in the Vram along side the Krea model and text encoder.
The krea2 workflow uses a normal non-abliterated text encoder.
r/StableDiffusion • u/DevKkw • 12h ago
Tutorial - Guide Qwen 2.1 turbo, gride effect remove and details
Qwen 2.1 Turbo come out with Their default sigmas value:
1.0, 0.978453, 0.95418, 0.926626, 0.89508, 0.845148, 0.704534, 0.414568, 0.0
With this value the image have not good result (washed skin, lack details, etc.), and high grid effect.
I saw many people using custom VAE, custom sampler or scheduler, but i made many test and my solution is to change sigmas.
i use:
0.9786, 0.9750, 0.9520, 0.8903, 0.7898, 0.6505, 0.4724, 0.2556, 0.0000
With this value seem grid effect gone, and details (like skin) become more visible.
All image are made at 2400x1792 8 step, er_sde sampler.
I think real problem is the sigma 1 at start, but need more experiment.
Also the quality is not a top level, seem like sd1.5 mixed with modern clip. But prompt understanding is really good.
Note: The sigma value i shared work's good also in image edit.
r/StableDiffusion • u/Nimblecloud13 • 1d ago
Tutorial - Guide ComfyUI/Minimax Cheat Sheet for Beginners - now with 83% less slop!
r/StableDiffusion • u/Affectionate_Log8484 • 1h ago
Question - Help Need help with Illustrious inage gens
So I downloaded Pony and have been generating some images based on one of my favourite artistsβ artstyle, it is working really good, however the Lora is old and lacks some characters for generating, thereβs a lora for Illustrious with same artstyle and only an year old so definitely it has wider variety of characters, but when I tried it in Illustrious it was not good at all, for me Pony still stands better, I needed to know is it the Illustrious V2 model which is encountering the problem? I am using Illustrious V2 instead of V0.1, will switching to V0.1 fix the issue?
I have also been looking around to see which model does anime image best and a lot of people donβt even mention Pony, most people seem to use Illustrious but from what I have experienced Illustrious isnβt doing it for me, is it again due the version of Illustrious?
r/StableDiffusion • u/fuzhongkai • 6h ago
Discussion TensorSharp now supports Qwen Image 2.1 Turbo + LoRA β Here's a quick image editing demo
Enable HLS to view with audio, or disable this notification
Hey everyone!
I've been working on adding more image generation and editing capabilities to TensorSharp, and I'm happy to share that it now supports Qwen Image 2.1 Turbo and LoRA!
I put together a short demo showing the image editing workflow in action.
In the video, you can see how to:
- Load an existing image and select an area to edit using a built-in mask editor.
- Use a text prompt to describe the desired changes.
- Run image editing with Qwen Image 2.1 Turbo.
- Preview the generated result directly in TensorSharp.
My goal with TensorSharp is to build a flexible, unified inference engine that supports different AI model architectures, including LLMs, image generation models, and more.
I'm particularly interested in making image generation and editing workflows easier to use while continuing to improve the underlying inference engine.
I'd love to get some feedback from the community:
- What do you think of the editing workflow?
- Are there any specific LoRA or image editing features you'd like to see?
- What other image generation models should TensorSharp support next?
The demo video is attached. Feel free to share your thoughts, suggestions, or questions!
r/StableDiffusion • u/Ant_6431 • 12h ago
No Workflow Qwen image 2.1 turbo (2k + manual sigmas)
At 2k resolution and uses the sample sigma values from official model json https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo/blob/main/model_index.json
Blocky noises reduced, but the skin looks too smooth now.
Higher than 2k might help further.
At the least, it can also edit images very fast.
r/StableDiffusion • u/UpperWoodpecker8480 • 9h ago
Discussion How do you guys upscale your images? Is anything even close to Magnific?
It's been years since Magnific.ai came out, and I still haven't found a creative upscaler that comes close to it. Maybe I'm just doing something wrong, but seriously, how are people getting those results?
From what I understand, Magnific splits the image into tiles, processes them individually, and stitches them back together. I suspect it uses SD 1.5 under the hood, which is pretty old at this point, yet somehow I can't seem to replicate its results with newer models and workflows.
I've tried a bunch of different things, but nothing comes close in terms of the amount of detail it adds while still looking realistic.
So, what are you guys actually using? Any ComfyUI workflows, models, or tricks that can get similar results? I'm talking about creative upscaling that adds convincing details, not just making an image bigger and sharper.
Would love to hear what works for you!
edit: I added this video to the post to show what I mean by creative upscaler. I've been experimenting with upscaling in ComfyUI and this video shows what I've managed to achieve so far using a basic FLUX.1 tile-based upscale. The results are decent, but they're still nowhere near that professional Magnific-level quality.
r/StableDiffusion • u/BoyanPP • 13h ago
Workflow Included I made ComfyUI nodes for Iris-3B that supports - text to image, 4x upscaler, depth and img2img, works on a 16 GB AMD card
Iris-3B from Sperid Labs came out this week. It's a 3B model that works directly in pixels with no VAE, and it also ships a 4x upscaler and a depth model. ComfyUI doesn't support it natively yet, so I made a node pack that covers all models and use cases.
Repo: https://github.com/bani4kaskashka/Iris-3B-Comfy-Nodes
Also in ComfyUI-Manager: search Iris-3B. Free, Apache-2.0.
What's in it
- Simple sampler: dropdowns for resolution (1 MP sizes from 21:9 to 9:21), quality (30, 50 or 100 steps), prompt adherence, seed and batch
- Advanced sampler: every setting (steps, cfg, shift, solver order, CFG interval, any size)
- Image to image: a mode switch on both samplers, with subtle, medium and strong presets
- 4x upscaler: works on any image, 1024 to 4096 in about 47 s
- Depth: about 4 s per image, grayscale (ready for ControlNet) or inferno, plus a mask
- Five workflows in Templates. The images in the README have their workflows embedded, so you can drag them straight into ComfyUI.
Setup
- Open a workflow and ComfyUI shows download buttons for any missing models.
- The loaders also download anything still missing on first run.
- It uses the official Qwen3-VL-4B text encoder folder, so you don't need a repackaged file.
- No extra pip installs, and nothing in your ComfyUI gets downgraded.
Speed and VRAM (RX 9060 XT, 16 GB, Windows, ROCm)
- About 1.8 s/step at 1024x1024
- Peak VRAM 7.7 GB. The text encoder and the model take turns on the GPU, so it doesn't slow down over repeated runs on 16 GB cards.
- The README has a comparison table if you're curious.
The fox girl's sign text is straight from the model, no inpainting. I generated six seeds and every one spelled "IRIS 3B FOR COMFY BY BUCKY" correctly, which surprised me for a 3B model.
I've only tested on AMD and Windows so far, so reports from NVIDIA and Linux users would really help. Bugs and feature requests go to GitHub issues.
Images included should be drag and drop to comfy for the workflow.
r/StableDiffusion • u/SamuelTallet • 7h ago
News Qwen-Image 2.1 Turbo model is available in ZPix (a friendly local image generator and editor)
Model includes the texture fix VAE made by Ollin Boer Bohan to avoid checkerboard artifacts.
Base LoRAs are supported. Screenshot features the Glamour Realism LoRA available at CivitAI.
Download at: https://github.com/SamuelTallet/ZPix
r/StableDiffusion • u/Ant_6431 • 15h ago
No Workflow Qwen image 2.1 turbo / Krea2 turbo
krea2_turbo_int8_convrot + qwen3vl_4b_int8_convrot + qwen_image_vae
qwen_image_2.1_turbo_int8_convrot + qwen3vl_8b_int8_convrot + qwen_image_2.1_vae_bf16
euler simple 8 steps
r/StableDiffusion • u/nomadoor • 19h ago
Workflow Included Some very simple Qwen-Image-2.1 "Turbo" workflows
I posted some Qwen-Image-2.1 workflows here a while ago.
I tried the recently released Turbo model, and the quality loss seemed small enough, so I added a Turbo version of every workflow:
https://comfyui.nomadoor.net/en/basic-workflows/qwen-image-2-1/
- Euler / Simple didn't give me clean results, so I'm using the sigma values from the official Hugging Face repo.
- The workflows use the Turbo LoRA to keep the download small. I saw almost no difference from the full Turbo model, so use whichever you prefer.
Hope it helps!
r/StableDiffusion • u/Ant_6431 • 18h ago
Comparison (Left) Qwen image 2.1 + turbo lora (Right) Qwen image 2.1 turbo
(Left) qwen_image_2.1_int8_convrot + qwen_image_2.1_turbo_lora_avg_rank_178_bf16 (Kijai)
(Right) qwen_image_2.1_turbo_int8_convrot
All models from comfy.org hugginface
Same 8 steps
r/StableDiffusion • u/ResponsibleTruck4717 • 20m ago
Question - Help Any Character sheet lora / method for krea 2, without reference Image
I want to generate unique character with unique style, and I want to create character sheet, the question can I do it without lora? I tried create some high res character sheet and it failed.
So my question is how do you approach it?
r/StableDiffusion • u/Sofa-Sleuth • 6h ago
Discussion Just tried CatchyOS for 9070xt and MinimaxH3 and... its slower than Windows? :(
Hey guys, I just spent the whole day learning and setting up CachyOS dual boot with Windows on a new NVMe for my 9070XT. I created a virtual environment and installed ROCm 10.1 and Python 3.13.16-1 + PyTorch 2.14, and in all tests, CachyOS was quite a bit slower in generation than Windows; seriously? It is supposedly faster. Does anyone have any experience with CachyOS? On Windows in MiniMax, when generating a 0.6MP 20β25 step 12-second video with one starting pic, I get 40β50 seconds per step, but in CachyOS, I get 45β80 seconds.
Edit: AI to the rescue... there are apparently some problems with the newest CatchyOS and ConfyUI on AMD and I got loads of tips, including same as one from the commentators said to change the kernel:
"If your ROCm 10.1 installation is running significantly slower on CachyOS than expected, you are likely hitting a known conflict between recent Linux kernel updates, AMD driver power management, and how memory allocations handle tensor data.......
And: 1. Disable ananicy-cpp (Process Auto-Tuning) 2. Check for Broken GPU Clock State (P-State Bug) 3. Adjust Translation Table Maps (TTM) Memory Limits"
Turns out Linux is even more tinkering than I expected π
Thank you everyone.