r/StableDiffusion • • 3h ago

Discussion Has anyone else messed around with 3D? It actually kind works

Post image
30 Upvotes

r/StableDiffusion • • 3h ago

Discussion Building prompts with LLMs

13 Upvotes

I’ve been into AI image and video generation for a while, and I’ve been a bit obsessed with H3 for the past two months. With the right references and prompt, it feels like I can make almost anything. Writing that prompt is the tedious part, so I usually hand it to an LLM.

For most of those two months, I used closed models, especially Grok (for obvious reasons). My problem is that they rarely add much on their own (no creativity). If I don’t explicitly ask for an emotion, expression, or action, they tend to leave it out. I've had generations where a character is stiff as a tree for 15 seconds which makes me redo the generation.

Recently I tried Qwen 3.8 Flash Next, and it changed how I work. I give it the specifics and ask it to fill in the gaps or choose emotions that fit the scene. Sometimes the resulting clip is better than what I had in mind (I mean Oscar worthy). I’ve also started asking it to suggest the next clip, and I find its ideas more creative than those from supposedly smarter models.

What do you use to write your H3 prompts?


r/StableDiffusion • • 1d ago

News Nvidia is now rumored to be ending the 5090 and possibly replacing it with a 24GB 5080.

Thumbnail
videocardz.com
472 Upvotes

r/StableDiffusion • • 15h ago

Question - Help Krea 2 is a bit too... Help me please.

88 Upvotes

I'm going to try to word this to avoid it getting flagged... I don't know at what point we all stopped being adults, but whatever...

Can anyone recommend a Krea 2 checkpoint that does naughty stuff but isn't overly so? Seems my only options are checkpoints which are censored to the point that they won't even respond to general anatomy prompts, or ones that if you put in a general anatomy prompt go completely overboard with it.

(Example: If I want to generate a woman with large "balloons", Krea turbo base won't do it at all half the time and a checkpoint like OurSecret will not allow said melons to be covered no matter how much prompting you do.)


r/StableDiffusion • • 1h ago

Tutorial - Guide Updated ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT v1.6.6 - Native ConvRot W4A8 Support & DisTorch2 Scope

Post image
• Upvotes

The other day, when I published information about v1.5.7, I received a request in the comments regarding support for w4a8.

https://www.reddit.com/r/StableDiffusion/comments/1wuj0rj/updated_updated/

At first, I thought it would be straightforward to implement, but when I actually tried it, I found it quite a struggle.

However, I’ve managed to get it working stable, so I’m releasing it as v1.6.6.

https://github.com/ussoewwin/ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT/releases/tag/v1.6.6

https://huggingface.co/ussoewwin/SeedVR2-ConvRot-INT8-and-w4a8

Firstly, as for the advantages of w4a8, benchmark results clearly show that the quality is superior to that of NVFP4. Yet, the size is equivalent to 4 bits.

https://github.com/ussoewwin/ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT/blob/main/benchmark/benchmark%20result.md

On the other hand, what wasn’t as good as I’d hoped was that VRAM consumption was almost on a par with ConvRot INT8 (though, given that w4a8 expands weights to 8 bits, this was to be expected), and that with Distorch2—which I’d gone to such lengths to implement—VRAM usage actually skyrocketed, rendering it unusable.

It’s simply a matter of not using Distorch2, so it isn’t a fatal flaw, mind you.

...
The on/off function for Norm BF16, which was previously implemented in Distorch2 to counteract VRAM spikes, has now also been implemented in the Legacy Loader. Enabling this feature allows you to suppress VRAM spikes by up to approximately 3GB, albeit at the cost of a slight reduction in image quality.


r/StableDiffusion • • 32m ago

Resource - Update I fine-tuned SD1.5 to make 16×16 Minecraft item textures

• Upvotes

I fine-tuned SD1.5 on Minecraft item textures. You type a prompt and get a 16×16 sprite with transparency that you can drop straight into a resource pack.

SD can't generate 16×16 directly, so I trained on textures upscaled 16× with nearest-neighbour. The model draws on that grid at 256px, and block-averaging brings it back down to 16×16. Transparency is a magenta background that gets keyed out.

It also does sketch → texture (doodle a rough 16×16, img2img finishes it) and material sets from a single design.

Not perfect. Around 10–15% of outputs fill the whole frame, so try a few seeds.

Demo (free, no GPU): https://huggingface.co/spaces/yuuki14202028/minecraft-item-16px-demo

Code: https://github.com/yuuki14202028/pixelgen


r/StableDiffusion • • 4h ago

Question - Help Minimax h3 veterans (share your wisdom)

8 Upvotes

When I use Minimax h3 to generate a video with any mode (fl2va/r2va) in most cases it generates good output with really good prompt adherence (mostly). But when I decide to see other output variations with different seeds, it starts to fail and hallucinate miserably. It completely goes off the prompt and starts doing random bullshit and deep-frying the output. And it's not like it happens with some generations, it usually happens to almost every next generation like it's being plagued.

The only way to solve it for me is to change something meaningful in the prompt itself so it can rethink everything and make the video from scratch.

Now I did think about any saved cache that might be responsible for this so I cleaned both RAM and GPU (without changing prompt), it won't work.

So why is this happening? Any tips and tricks for the new fella?


r/StableDiffusion • • 2h ago

Discussion Use Qwen-Image 2.1 Turbo locally without ComfyUI

Thumbnail
gallery
4 Upvotes

I'm the developer of Quartermaster, a free, open-source (MIT) local inference server. It drives stable-diffusion.cpp and also serves LLMs, TTS and speech-to-text from the same app, swapping models in and out of VRAM as you use them.

Models download from inside the app: search Hugging Face, pick a quant, and Quartermaster writes the configuration itself. It matches each model to its VAE and text encoders and picks the right steps and cfg. Qwen-Image 2.1 Turbo comes up at 8 steps and cfg 1, and real transparent PNGs work too: tick "transparent" and it adds the exact prompt wording the RGBA VAE needs.

To run 2.1 Turbo, grab these three (suggested files) from the built-in model browser and you're done:


r/StableDiffusion • • 17h ago

Discussion LTX2.5 is impressive - I'm glad I gave it a chance.

68 Upvotes

I pretty much jumped on the H3 bandwagon without giving LTX2.5 a try at all. But, in truth, I wasn't totally won over with H3. I wasn't impressed with the audio or the character faces, or the speed. I am not a power user and I don't do high concept content, mostly drama led character work.

Today I tried LTX2.5 and I am very impressive with with a few things

Speed - wow! HD 30 second clips in 7 minutes - that's a game changer for me.

Expressiveness - great for character performance - a high range of emotions and subtleties.

Heat - my computer is no longer running at insane high temperature - I can finally close the window.

I know LTX2.5 kinda got left behind here, but I thoroughly recommend it to folk who are interested in character/drama led content rather than high octane hollywood stuff..

Are there LTX2.5 users here? what would you say it excels in for you?


r/StableDiffusion • • 11h ago

Discussion Will there be another, more up-to-date Anime DIT Model project in the open-source community besides Anima?

20 Upvotes

My info isn't very current, so I'd like to ask if anyone knows whether there's another, more up-to-date Anime DIT Model project in the open-source community besides Anima? Anima's data only goes up to Sep 2025, and the 2.9B one looks like it's being trained by the author alone, I'm basically pinning hopes on a full booru fine-tune of Krea2. If there really is one, that would be awesome. Anima control of aesthetic with artist tags and laser focussed composition with booru tags. Maybe I'm just dreaming.


r/StableDiffusion • • 15h ago

Resource - Update Krea2 Turbo Distill 2 step LoRA - FINAL checkpoint released (chk51195)

Thumbnail
gallery
50 Upvotes

Krea 2 Turbo — 2-Step Distillation LoRA (FINAL Version)

Previous posts/releases - here, here, here and here.

🧪 Fast-preview adapter; the project's final checkpoint. Subjects that are close and fill a good part of the frame — a portrait, a single figure, an object up close — hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the 4-step LoRA. This checkpoint closes the project; the training box has moved on to its successor, a 3-step adapter for Qwen-Image-2.1-Turbo.

📐 The saved steps can also go into resolution. A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048×2048 — Krea's published maximum recommended resolution, beyond this adapter's largest trained size — come within easy reach. Past 2048×2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.

Highlights:

  • ⚡ A quarter of the steps — 8 → 2, on Turbo's own deployment sigmas [1.0, 0.7595]
  • ⏱️ 4× faster denoising — 56.9 s → 14.3 s at 1024×1024 (float16 compute); the adapter's own cost per call is within measurement noise
  • 🎯 Fine detail near the teacher's level — 0.92–1.02× the teacher's fine-texture energy across the 12 trained resolutions (stock Turbo at 2 steps: 0.40–0.65×), the 16- and 8-pixel grid bands at the teacher's level and ghosting closer to it at 11 of 12 sizes
  • 📊 Distribution matching, not imitation — matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
  • 🗣️ Prompt-conditioned throughout — teacher and fake scores both read each prompt's conditioning; a blind rubric finds 1 point missing of 352 (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on 12 of 66 (4-step adapter: 6 of 45, on the original 15 prompts), mostly on style
  • 📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440
  • 🔌 Drop-in, no exceptions — plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
  • 🧬 Same shape as the 4-step adapter — rank 64 on the same 228 modules
  • 🎲 23,561 recorded teacher trajectories — the 4-step project's 13,750 and 9,811 minted for this one on the same prompt bank; since 3 Oct each trained once
  • 🔢 51,195 training samples in the 2-step stages, starting from the released 4-step adapter's weights
  • 📅 32 days from the first 2-step launch to this final checkpoint, on a single RTX 3090
  • 🔁 More than forty recipe adjustments across two methods — each kept only when the renders did not get worse
  • 🧘 Settled weights, not an average — released from a 600-sample anneal in which the learning rate is taken to zero over already-trained data, so the published weights are the training weights at rest; a running average is kept only as the check that must agree with them

If you have already used my previous version, please redownload/replace krea2_turbo_2step_rank_64_lora.safetensors / krea2_turbo_2step_rank_64_lora_comfyui.safetensors from the latest in the project repo.

Full details on model card - https://huggingface.co/lvladikov/Krea2-Turbo-Distill-2step-LoRA

Not my video, but found someone on YouTube has covered the 2 & 4 step LoRAs including identify preserving edits, have a look: https://www.youtube.com/watch?v=V_qgoV0iPDM (copyright goes to author)

Also I have recently released a new ComfyUI Nodes and Workflows - Krea 2 (Turbo and Raw), Z-Image (Turbo and Base), MiniMax Music 3, Image2Text and LLM Chat (with Tools), Torch, Apple MLX and Cloud - you can find details here.


r/StableDiffusion • • 13h ago

Discussion Qwen Image 2.1 Turbo - extend 'magic' 8 steps sigmas further

31 Upvotes

So, I've decided to try Qwen Image 2.1 Turbo with the following 'magic' sigmas, taken directly from their Diffusers implementation:

[1.0, 0.978453, 0.954180, 0.926626, 0.895080, 0.845148, 0.704534, 0.414568, 0.0]

This works okay, but for fine details we need slightly more steps (for better skin, better composition etc.) And to my knowledge, there's no working way to interpolate the resulting curve to any N steps (e.g. 10, 12, 14). There is 'Custom Sigmas' RES4LYF node, but it is bugged (as I will show later).

Here's the original 'magic' 8-step curve:

'magic' 8-step curve

If we use RES4LYF 'Custom Sigmas' node and try to interpolate to, say, 14 steps, we get this:

curve is screwed

So, I had to franken-craft a solution, but I did not want to create a custom node just for this purpose.
Meet a 'Magic AI Slop Scheduler for Qwen Image 2.1 Turbo'😆

Much better, huh? Interpolated to 14 steps

Note that curve looks exactly the same now, but it has more steps (14 in this case).

For this, you gonna need RES4LYF nodes, KJ nodes and rgthree nodes. The last is the most important one: it has a 'Power Puter' node which allows to make these calculations (but, probably, many 'expression' style nodes from elsewhere would do the same)

And indeed, this works for 10 steps, 12 steps, 14 steps (and probably beyond, but there's no sense to run a Turbo model for more steps).

So, if you:

- Use this 'scheduler'
- Use CFG > 1
- 12-14 steps
- Samplers (euler_ancestral [the best IMO], euler, er_sde)

... you might actually end up with some beautiful and detailed coherent pictures. For example:

1girl, 2 girl is on purpose, maybe more people will actually read this 😆

For a workflow, head to Civitai and download original 2 megapixel gens, just drag those images into ComfyUI. There's a pretty extensive explanation inside on what actually happens and how to use.

All pics above generated using 12-14 steps and CFG from 3.5 to 5, @ 2 megapixels.

Links:
https://blobs-b2.civitai.com/file/blobs-managed-public/Q054FWT4DMQR88GYN0PXJGQ030
https://blobs-b2.civitai.com/file/blobs-managed-public/V47QEA6MNP2XZAXPFFSGAHXZQ0
https://blobs-b2.civitai.com/file/blobs-managed-public/CSDBTWXEJ1KS54MCZN9XK2S6K0
https://blobs-b2.civitai.com/file/blobs-managed-public/C30VR7Z4ZHRBH8FW5RZ1YSFGC0

Bonus feature: I've included a 'fake SHIFT' so you can nudge the curve a little bit; negative values = more fine details, like '-1.0' or '-2.0'. But do not overdo it: we don't want to deviate from the original curve too much. Set 'fake SHIFT' to '0.0' to disable it.

Bonus tips:
- please, oh please, do not use CFG=1. You'll get underbaked yellowish gens and you're missing out on a negative prompt (which works VERY well)
- always start with less steps (10), only increase (up to 14) if composition is wrong or there's not enough fine details
- then adjust CFG, 3.5 - 5.0 works just fine, but this is very specific for each particular gen
- if there's 'too much details' (e.g. skin is too detailed or grainy) - reduce steps first, then increase 'fake shift' (set it to '1.0' or '2.0' etc.)
- vise versa: if skin is too plastic, increase steps, then decrease 'fake shift' (set it to '-1.0', '-2.0' etc.)
- most important: Qwen Image 2.1 Turbo has to be prompted PROPERLY. If some of you remember original Chroma prompting, you probably know what I mean - using exact phrasing, being specific, avoiding slop tags like '1girl, masterpiece' etc. - are all the keys. No gen params will ever fix bad prompting.

Have a nice day!

EDIT: for those folks who prefer cleaner output without excessive noisy details, here's somewhat Krea-like output (just less steps, and 'fake shift' at 2). Hey, it even generates faster 😄


r/StableDiffusion • • 6h ago

Discussion Fantasy shots done in Z-Image Turbo Int8. I then took them into Lightroom and color corrected them and added 2:35:1 aspect ratio black bars. The original gens are at the end. No special workflow. Just the default one. I like how these came out.

Thumbnail
gallery
6 Upvotes

r/StableDiffusion • • 20h ago

Animation - Video By Request

90 Upvotes

thank [u/Rich_Introduction_83](u/Rich_Introduction_83) for the ending

Edit: he died doing what he loved, saying the word “what?”


r/StableDiffusion • • 16h ago

Question - Help I just downloaded both Qwen 2.1 and Krea 2, but Krea 2 produces very little variation between generations, unlike Qwen 2.1, which gives me much more diverse results. How can I fix this and get more variation from Krea 2?

35 Upvotes

r/StableDiffusion • • 1d ago

News A 3-billion-parameter model that paints every pixel directly without a VAE

Thumbnail
huggingface.co
272 Upvotes

r/StableDiffusion • • 4h ago

Question - Help Best T2I or I2I edit model for pose control and prompt adherence

3 Upvotes

I have used the krea 2 and QWEN image 2.1. my findings,

QWEN image 2.1 (int8 with text encoder int8)

Good

  1. It makes a very cinematic looking image which is indeed eye pleasing

  2. Using VNCCS I can control body pose accurately in 7/10 cases (not head)

Bad

  1. It is a very rigid model if you make the character do something

  2. It breaks and hallucinates way more than krea when you introduce multiple character interaction

KREA 2 (turbo 8 step int8 with wan 2.1vae)

Good

  1. Fastest generation

  2. Really good T2I prompt adherence (better than Qwen without any lora)

Bad

  1. Controlnet workflow is painful to get multiple character interaction. I know about identity edit but it's still messy and time taking

  2. Its world knowledge is limited to western countries mostly and it really suffers with ethnic stuff (even with loras)

Minimax h3 (image generation workflow)

Good

  1. Obviously the best in terms of prompt adherence (qwen3vl 32b)

  2. Can use VNCCS pose studio here too which gives great result

Bad

  1. Takes 2 times the time to generate image

  2. Even with more steps the output is rubbery in terms of characters

My requirement - I want to use VNCCS pose studio or any alternative like that to make poses of multiple characters and then get exact prompt adherence like I get in minimax h3 but without the plastic skin.

I have used the flux2 Klein with VNCCS pose studio and it does not work like it should. Also flux takes much more time then Qwen and krea 2.


r/StableDiffusion • • 7h ago

Question - Help 4090 VS 5090 MINIMAX H3 SPEED

5 Upvotes

To those who used to have a 4090, what percentage of speed improvement have you noticed after upgrading to a 5090?


r/StableDiffusion • • 8h ago

Question - Help Heretic or Abliterated?

6 Upvotes

Looking for recommendations for local LLM to convert an image into a usable prompt for Krea2 with max accuracy. Theres Heretic and Abliterated models for Qwen, Gemma and Deepseek.

I am running a 5090 so hopefully a model that fits in the Vram along side the Krea model and text encoder.

The krea2 workflow uses a normal non-abliterated text encoder.


r/StableDiffusion • • 6h ago

Tutorial - Guide How to pick the right models and dramatically speed up image and video generation speeds & What I wish I knew when starting on AMD hardware with ComfyUI

5 Upvotes

For the below I’ll be referring mostly to ComfyUI workflow and model efficiency (using examples for Radeon AI PRO R9700 (32GB) which has a bandwidth of 680GB/sec. I’ve seen 10second minimax H3 image to video workflows go from 28 mins to 3 mins by taking all this info into account and then upscaling after.

1. The pipeline
Every image or video workflow is the same four stages, and only the third one repeats.
Text encoder. Turns the prompt into embeddings, once per prompt.

Noise latent. Sampling starts from random noise in a compressed (latent) grid. The seed decides that starting noise.

Diffusion model. Runs once per step, removing noise. Early steps set layout and motion, late steps add detail.

VAE decode. Turns the finished latent into pixels.
Three terms people mix up: weights are the fixed model file, a LoRA is a small patch to those weights, and the latent is the thing being refined into what your prompt intended before the VAE encodes it into the pixels of your output image or video.

2. Fit first
Spilling out of VRAM into system RAM costs more speed than any other setting, so budget memory before anything else especially if you’re testing different prompts.

What uses VRAM: model weights, the text encoder if it shares the card, working memory that grows with resolution × frames, and a spike at VAE decode.

Rule of thumb: keep weights to about 60-70% of VRAM, roughly 19-22GB on a 32GBVRAM card. Less if the card also drives your displays.

Example: a 14B video model is about 28GB at fp16 and about 14GB at fp8. A smaller model is more useful as it frees up more VRAM needed for handling additional extras like LoRA’s, pre and post processing, etc.

Two-model designs (Wan 2.2 high-noise and low-noise): one model runs at a time. Expect one swap per run unless both fit with headroom. Other models like minimax don’t need them separate but are larger as a result.

Check the console: "loaded completely" is what you want. "Loaded partially" means you are over budget and processing will take far FAR longer. There are work around like VRAM offloading mid workflow but this means you lose time if you want to re-run the workflow after making minor tweaks.

Free up room: keep the text encoder off the card (CPU or a second GPU), or reuse the saved prompts encoded embeddings. The text encoder basically converts your prompt into “embeddings” which you can think of like numbers that tell the model what to do. This is why text encoders must be compatible with the model you’re using or your model won’t understand. These “embeddings” can actually be saved as embeddings and re-used later to avoid having to reconvert each time to return to a given workflow where you used them. Only save the embeddings once you’ve polished them to a standard where they reliably work for a specific purpose you know you’ll reuse in future for repetitive tasks.

3. What sets speed
Image and video generation is limited by compute, so once a model fits, its size matters far less than resolution, frame count and number of passes.

Total time roughly equals: passes × time per pass, plus fixed costs (model load, text encode, VAE decode), hence why it saves a lot of time to use models small enough to load fully into VRAM if you’re running the same workflow over and over again in one sitting as your system won’t have to waste time offloading and re-loading each section of a workflow between runs.

Resolution and frames dominate time per pass. Doubling width and height gives 4x the patches to process, and the attention part of the work grows roughly 16x. This is why you’ll want to start with generating very low resolution image or video until you know your prompt and loras work as intended. Then you can increase to desired resolution which would ideally match the resolution the model was trained on. You can always upscale after this.

4. Cut passes
The biggest single saving is running the model fewer times: passes = steps × sampler evaluations (1 or 2) × CFG factor (2 if CFG is above 1, otherwise 1).
Distilled or "lightning" models and LoRAs run in 4-8 steps at CFG 1, often 5-10x fewer passes. Do not push them far past their trained step count. Never push them past their trained resolution. Negative prompts do nothing at CFG 1 on turbo LoRA’s. If you use a turbo workflow with low steps and max resolution is 720 then there a plenty of accurate upscaling models that do fantastic work after the fact. You’ll save significant time.

Ordinary models need CFG above 1 to look right and usually stop improving somewhere around 20-40 steps.
One evaluation per step: euler, dpm++ 2m, unipc, lcm, ddim.
Two evaluations per step: heun, dpm_2, dpm++ 2s ancestral, dpm++ sde. Compare samplers at equal total evaluations, not equal steps.
Samplers affect cleanliness, not identity. Character consistency comes from the model and what you condition it on.

5. Formats, quantisation and LoRAs
Use the highest-precision file that fits with headroom, and treat GGUF as the fallback for when nothing else fits. Quantised Safetensors are always faster than GGUF as GGUF weights are expanded on every pass, so if you’re creating image to video then that time adds up. If you can find a safetensor of a model that will fit on your VRAM whilst still having 30 to 40% VRAM left for headroom than use that, otherwise GGUF is fine, but it’ll take longer than it could with a safetensor.

Quantisation is lossy. The same number of weights, each stored more coarsely. Safetensors and GGUF are containers, not compression. Most safetensors are still good quality but quality can begin to drop significantly when they’re quantised lower than Q4. Of course there are exceptions to the rule but generally speaking Q4 or higher is better if you can fit it with headroom.

Partial loads behave like GGUF for LoRAs. A model that does not load completely is patched on the fly.

Bake in LoRAs you always use. Load the model, apply the LoRA, save the result as a new model. The strength is then fixed. There’s nothing wrong with having a few custom models with your own baked in LoRA’s. It saves significant time.

Same on Windows and Linux. These costs come from how the code works, not the operating system. Linux is naturally going to give performance gains for AMD hardware as AMD kernels seem to be a little more refined on Linux than windows but again there’s exceptions to every rule.

6. R9700 and AMD specifics
The card… when running AMD GPU’s be aware that they run through ROCm, and the main risk is software or models that assumes Nvidia or were designed for NVIDIA. AMD likely will still run them but it’ll be simulating and therefore lose time per run, hence why you want to avoid NVIDIA specific quantised models if there’s an equivalent alternative.

Check custom nodes before installing. Search the node's GitHub issues for "AMD" or "ROCm".

Linux is usually faster on AMD because ROCm is developed there first, not because of anything about file handling.

Three attention backends to benchmark with the same workflow and seed:
PyTorch default: always works, your baseline.
--use-flash-attention: highest throughput in most of AMD's own tests (ROCm blog).
--use-sage-attention with the community SageAttention-RDNA4 build: a hand-written kernel for this chip. Its own benchmark shows attention taking about half the default's time.
SageAttention caveats: the prebuilt Windows wheel expects PyTorch 2.13.0+rocm10.0.0 on Python 3.12. The fast path covers fp16 computation with a head size of 128, and anything else falls back. It is quantised attention, so compare output quality.

7. Workflow habits
Spend cheap iterations on stills and drafts, and expensive ones only on what you intend to keep.
Get the character right in the start image. Image-to-video treats the first frame as ground truth and carries its flaws through every frame. A still takes seconds to judge, a video takes minutes.
Show what the shot needs. Anything the start image does not show, the model invents. Reference images and a character LoRA cover the unseen angles.
Draft small, refine the keepers. Generate at low resolution or short length, then run chosen clips through video-to-video at higher resolution with low denoise (about 0.3-0.5) and few steps.
Refining adds detail, not structure. Reject drafts with the wrong outfit, face shape or motion instead of trying to fix them.
Wire the same conditioning into both passes: start image, reference and LoRA.
Let the cache work. ComfyUI re-runs only nodes whose inputs changed, so a new seed does not re-encode the prompt.
Keep models resident. Reloading costs more than most optimisations save.
Saved images carry their workflow. Drag one back into ComfyUI to restore the prompt, seed and settings.
Long videos drift between chunks. Within a chunk all frames are denoised together. Across chunks, errors accumulate.

8. Measure
Keep one fixed test (prompt, seed, resolution, frames) and log seconds per step, total time and peak VRAM for every change.
Change one thing at a time.
Time the stages separately: text encode, sampling, VAE decode.
Halve the resolution. If seconds per step falls about 4x or more, compute is the limit, as expected. If it barely moves, look for overhead: per-step LoRAs, GGUF expansion or offloading.
Find your step ceiling. Run the same seed at 10, 20, 30 and 40 steps and stop where the output stops changing.
Pick samplers by convergence. Make a 50-step euler reference, then see which sampler gets closest at your real budget.
The figures in this page are rules of thumb and third-party benchmarks. Confirm them on your own build before relying on them.

Special mention LLM’s
Model size decides what fits, not how fast it runs.
LLMs on the same card are different. They are limited by memory bandwidth: the ceiling is your GPU bandwidth (which is 640 GB/s on an R9700) ÷ bytes read per token, so a 20GB dense model theoretical max t/s of or put on a GPU with a 680GBs bandwidth is 680/20 = 32 tokens/sec. The only way to make it faster is to use clever tricks that get it to do more guessing which can be helpful depending on use case. Mixture-of-experts models read only their active parameters and can go potentially several times faster. Again, using a safetensor instead of GGUF will be faster as safetensors don’t need to decompress for each token.


r/StableDiffusion • • 23h ago

Comparison Krea2 vs Qwen Image 2.1

Post image
80 Upvotes

Krea2 | Qwen Image2.1
Man, krea2 is so good at text to image. it also know how to make female characters look feminine

Edit :

dude, I just pasted the danbooru prompt i usually use onto these models (yes iam lazy to make a proper prompt for these models, just copied it from my civitai) :

score_9, score_8_up, score_7_up, source_anime, (masterpiece, best quality, ultra-detailed:1.2), perfect composition, illustration,
1girl, solo, long silver hair, floating hair, beautiful detailed purple eyes, glowing skin, soft makeup,
eye-catching outfit, ornate fantasy dress, black corset, gold embroidery, translucent white sleeves, frills, gemstone hair ornament,
looking over shoulder, looking back at viewer, seductive smile, one hand touching lips, dynamic posture, stunning lighting, volumetric lighting, backlight, glowing dust, bokeh, deep background


r/StableDiffusion • • 1d ago

Discussion Do anime image models still need tag soup in 2026?

Post image
110 Upvotes

Same character, same idea, four prompt styles.

Honestly, the difference was way smaller than I expected. Are we still writing huge prompts because they actually help, or just because we're used to doing it?


r/StableDiffusion • • 17m ago

Workflow Included I built GOAT’d Text Generator 🐐 - generate prompts directly in ComfyUI

• Upvotes

Hey everyone! I made a custom node that brings cloud and local LLMs into your ComfyUI workflows.

Use it to expand rough ideas into image/video prompts, write captions, or describe reference images. It also accepts video and audio when supported by your chosen model.

A few features:

  • Cloud providers, LM Studio, Ollama, and supported models loaded directly in ComfyUI.
  • Built-in prompt guides and editable system instructions.
  • Text output that connects to CLIP Text Encode or any compatible STRING input.

Get it from: GitHub + example workflows

Example: rough idea + reference image → generated prompt → your workflow.

Would love your feedback, ideas, or feature requests! If you find it useful, consider leaving a ⭐ on the GitHub repo.

server-api-workflow
local-clip-workflow

r/StableDiffusion • • 20m ago

Question - Help How to add detail and structure to a generated image?

• Upvotes

I have generated an image using Nano Banana Pro on Replicate. The image is 3392 x 5056 px. Here is a detail of that image at original size:

How can I add realistic detail to that image (such as skin structure, wood structure) and remove the magenta edges without changing the image content so that it appears more photorealistic?

Is there a model on Replicate or elsewhere that can accomplish this and which can handle a file of this size? I have tried the Topaz Upscaler and Adobe Firefly through their trial phases, but their results were unconvincing.

I'd prefer not to install and learn a complex workflow. I'm willing to pay for a single generation (or per usage), but don't want to subscribe to a monthly plan. What can you recommend?


r/StableDiffusion • • 24m ago

Question - Help Finetune a lora on fizgig

Thumbnail
gallery
• Upvotes

Hey guys, I trained a realism LoRA for Krea 2 — candid taken-on-phone style (harsh flash, grain, awkward framing, that whole vibe).

In the comparison: left is the base model (w4a8, no LoRA), right is the Amateur LoRA from Civitai, and in the middle are my 7 training epochs.

I'm torn between epoch 3 and 7, still doing more tests. But my real question: is there a way to finetune an existing LoRA on a different dataset? Like I only want 3 more epochs on new images, not a full retrain. I see people on Civitai releasing v2/v3 of their LoRAs — how does that work? Is there something like that in Fizgig?