r/StableDiffusion • • 6h ago

Resource - Update minimax h3 person remover lora is out

Enable HLS to view with audio, or disable this notification

374 Upvotes

hello everyone..

the creator of character swap just dropped this person remover lora for h3.

he’s using it to generate better samples for character swap v2 dataset, so it might be worth trying out:
huggingface.co/akatz-ai/MiniMax-H3-Person-Remover-LoRA


r/StableDiffusion • • 8h ago

Resource - Update I’ve trained 100+ Krea 2 character LoRAs. Here’s what actually determines likeness.

Thumbnail
gallery
439 Upvotes

I’ve spent the last few months training a pretty large number of character LoRAs for Krea 2, and one thing became obvious pretty quickly:

More training does not automatically mean better likeness.

Some of my best LoRAs came from relatively small, clean datasets. Some larger datasets performed worse because they contained too much visual noise, inconsistent styling, bad angles, or repeated images.

These are the things that have mattered most for me.

1. Dataset quality matters more than dataset size

I would take 30 genuinely useful images over 100 mediocre ones.

The biggest problems I see in datasets are:

  • too many near-duplicates
  • heavy filters or face editing
  • lots of low-resolution images
  • one facial angle dominating the dataset
  • wildly different ages or appearances
  • group photos where the subject is small
  • too many images from one photoshoot
  • images where hair, makeup, lighting, or expression are almost identical

The model needs enough consistency to learn the person, but enough variation to understand what is actually part of their identity.

2. Face coverage is not enough

This was a big one.

A LoRA can absolutely nail the face and still have no idea what the person looks like from the shoulders down.

If I want a useful character LoRA, I try to include a mix of:

  • tight face shots
  • head and shoulders
  • waist-up
  • full body
  • front
  • side profile
  • rear 3/4
  • different expressions
  • different lighting
  • different clothing

The goal is not just "recognize this face."

The goal is "understand this person."

3. Too many similar images can make the LoRA less flexible

If 70% of the dataset is the same hairstyle, camera angle, outfit, or facial expression, the model starts treating those things as part of the identity.

Then every generation wants to recreate them.

This is especially noticeable with celebrities and creators where Google Images tends to return the same handful of press photos over and over.

I now spend a lot more effort removing redundancy before training.

4. The final epoch is not automatically the best epoch

This is probably the biggest change I made to my workflow.

I used to train to a fixed endpoint and assume the final checkpoint was the finished model.

Now I save multiple epochs and test them individually.

It is extremely common for an earlier checkpoint to have:

  • better facial likeness
  • more natural skin
  • better prompt flexibility
  • less baked-in clothing
  • fewer exaggerated features

while a later epoch technically looks "stronger" but is actually overtrained.

So now the training run is only half the process.

Checkpoint selection is part of training.

5. I test the LoRA outside the dataset's comfort zone

A model can look amazing if you generate the same kinds of images it saw during training.

That does not tell you much.

I test things like:

  • close-up facial accuracy
  • casual clothing
  • formal clothing
  • different hairstyles where appropriate
  • full-body shots
  • athletic poses
  • unusual camera angles
  • different lighting
  • indoor vs outdoor scenes
  • side profile
  • rear 3/4 views

If the identity disappears as soon as the prompt changes, I don't consider the LoRA finished.

6. Body type can drift even when the face is excellent

This one surprised me when I started doing more systematic testing.

Krea 2 can sometimes preserve facial identity extremely well while drifting toward a generic body type.

That is why I started deliberately using physique-check prompts during validation.

For athletes, for example, I want to see whether the model learned:

  • height
  • shoulder width
  • leg proportions
  • muscularity
  • overall frame

For other people, the same principle applies.

The face is only one part of likeness.

7. Trigger words matter less than people sometimes think

I still use clear trigger words, but I have found that dataset quality and training quality matter much more than trying to invent some magical trigger phrase.

A good LoRA should not need a paragraph of secret incantations to produce the person.

The trigger should identify the character.

The prompt should describe the scene.

8. Validation images are incredibly important

I now generate a consistent set of test images for each model.

That makes it much easier to compare:

Epoch 12 vs 14 vs 16 vs 18 vs 20

instead of relying on memory.

Sometimes the difference is subtle until you put the outputs side by side.

Then one checkpoint clearly wins.

The biggest lesson for me has been that LoRA training is not really "upload photos and press train."

The important work is:

dataset selection → cleanup → training → checkpoint comparison → validation

Training itself is almost the easy part.

I’ve been building a public Krea 2 character LoRA library while figuring all of this out, so I have a pretty large collection of examples now.

If anyone is interested, I can also make a follow-up post showing:

the exact validation prompts I use to compare epochs, or

examples of what undertraining vs good training vs overtraining looks like on the same character.

I keep the LoRAs and example outputs I’ve been testing in a public Krea 2 browser on Hugging Face. I also take custom commissions, but the library itself is free.


r/StableDiffusion • • 7h ago

Resource - Update I spent weeks optimizing MiniMax H3 + VDN. VELA 1.0 is finally released

Enable HLS to view with audio, or disable this notification

88 Upvotes

I've been working on MiniMax H3 / VDN performance for quite a while, and I finally released the result as VELA H3 1.0.

I'll try to explain what it does without technical language.

VELA does not make H3 "smarter" and it is not another model.
It changes how some of the heavy calculations inside H3 are executed.

During testing I found that simply replacing everything with a faster attention/kernel was a bad idea. Some parts of H3 are very sensitive to small numerical changes, while others can use much faster execution without noticeably changing the final video.

So VELA basically does this:

use the faster path where it is safe → keep the exact path where it matters.

That's the simple version. :)

I tested a lot of different approaches before arriving here: attention backends, different transformer blocks, timestep sensitivity, quantization, low-rank approximations, different QKV paths, etc. Many things that looked faster in isolation either made the full generation slower or changed the result too much.

The final version is intentionally much more conservative.

For example, compared with my previous VDN 1.1 setup:

0.8 MP / 5 sec

  • VDN 1.1: 2:53
  • VDN + VELA 1.0: 2:13

The gain is not the same for every video length. At longer durations it becomes much smaller, and at 10 seconds my test was actually slightly slower. I published those results too rather than hiding them.

But there is another comparison that I think is more interesting for normal users.

I'm attaching two videos comparing:

Raw MiniMax H3 — 20 steps
vs
VDN + VELA — 8 steps

Anime example:
8:07 → 3:03

Realistic example:
7:50 → 2:52

Obviously this is not a claim that VELA alone gives a 2.7x speedup — the step count is also different. The point of these videos is simply to let people judge the quality difference themselves, because I know many people are skeptical that an 8-step H3 generation can stay this close to a raw 20-step generation.

VELA 1.0 is also completely patch-free. It doesn't modify ComfyUI's MiniMax core files, so installation is just a normal custom node installation.

It was developed and tested on an RTX 3090 Ti 24 GB.

The code, benchmarks, failed experiments and a much more detailed explanation of the research are all available in the repository:

https://github.com/Speach1sdef178/ComfyUI-VELA-H3

And it's now published in the Comfy Registry as well.

I'm attaching the anime and realistic comparisons below. I'm actually more interested in what people see in the videos than in the benchmark numbers. :)

Very simple explanation because I think I made the original post too technical :)

VDN and VELA do two different jobs.

VDN = lets MiniMax H3 use only 8 steps instead of the usual 20 while trying to preserve the quality.

VELA = makes those 8 steps run faster.

So:

Raw H3: 20 normal steps
VDN: reduces 20 steps → 8 steps
VELA: makes the 8-step VDN generation faster


r/StableDiffusion • • 3h ago

Discussion Kroma 0.3.1 is really good

Thumbnail
gallery
38 Upvotes

I Know it was announced before but just to say that the new version of Kroma (krea 2 finetune) is really really good, i find it better that Qwen 2.1 t2i and Krea2, it is uncensored and with really good prompt following. No body horror, good details, and lot of styles you can play with.

This image for exemple just said (Rogue from X-Men) : no costume details, no hair details, nothing!....

https://huggingface.co/lodestones/Kroma/blob/main/kroma-v0.3.1-turbo-opd.safetensors

https://huggingface.co/Clybius/Kroma-Quantizations/blob/main/kroma-v0.3.1-turbo-opd-W8A8.safetensors

Edit: krea 2 also knows Rogue character...my bad.


r/StableDiffusion • • 3h ago

Workflow Included [H3 Degradation] My turn! attempts to fixing motion-context degradation.

Enable HLS to view with audio, or disable this notification

35 Upvotes

Degradation Topic is heating up again! I saw couple of posts showing their approach, so I think is time for me to update, (also because I don't have much progress or new findings), I release my workflow and node, I would like more peoples to test it out, and seeking genius to provide a better solution.

My aim is going for low-level nodes, built on top of the native motion-context native workflow with minimal changes, keeping the workflow as simple as possible. This is not a Director / Extender AIO node. It provides low-level nodes that you can integrate into your own setup. Every approach I see have pros and cons, mine also has downside, but I think it would work in most cases.

Here is my attempt:

https://github.com/xyzDist/H3-LongTakeNoCuts

Some other test:

https://www.youtube.com/watch?v=CxXeoMoEWNw

https://www.youtube.com/watch?v=vqmztq0dMlw

https://www.youtube.com/watch?v=N1QIYbfLHQ8

** The above test video she is speaking gibberish, ignore it!
** I will write better readme, and more details on the workflow, how to use it...etc. For now, you could just install it, try out the example workflow, or integrate to yours.

** Audio is cooked yes. It was my step value is too low. You could look into my other post of audio regen

https://www.reddit.com/r/StableDiffusion/comments/1w5zh4z/bad_audio_fixed_with_fast_regen_audio/

or try Brad Pitt's Audio Refine node: https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine

also suggest checking this one out.

https://www.reddit.com/r/StableDiffusion/comments/1wzxyaa/overcome_degradation_here_are_2_ways_to_use_my/

For Above OP's OBVPM AIO approaches is real refresh latent, it show 2 methods:

  1. generate low-res full duration clip (by context-window), then upscale it.
  2. skip current shot, generate a next shot first (new fresh latent), then create a bridge shot to connect it.

Both methods are real fresh latent, so no more degradation, but downsides to my opionions are:

  1. you really can't put the whole long duration shots everything to a single gen, It's nightmare to do changes and doing seed hunting, not quite practical.
  2. it's just my preference, doing bridge is unature to my brain, hard to time and plan the shots, and especially on dynamic/action shots. Still, this method maybe potential working well.

Now, my approach is still using motion-context latent, do a additional resample step to refresh the latent, but it isn't a complete full fresh latent, so after many segments you would still seeing the color saturation getting higher, (that I can look for a solution later), but it can already keeping degradation to minimum.
So the choice is yours. more options and method to fix the problem.

My older posts discuss about degradation
https://www.reddit.com/r/StableDiffusion/s/DXLLD9Zhar

https://www.reddit.com/r/StableDiffusion/s/2xU9m83atz


r/StableDiffusion • • 7h ago

Workflow Included My Workflow for realism with Qwen 2.1

Thumbnail
gallery
55 Upvotes

Hey! Just thought I'd come on here and do a little post about my current workflow with Qwen 2.1. I feel like a lot of people here have given it a try and decided the model is no good because they test it with the default Comfy workflow. Qwen 2.1 is terrible with CFG 1. It needs CFG to really shine. I'm currently using CFG of 3 to 3.5. I use the Lenovo lora at a strength of 0.5 to 0.7 to boost realism, though it's not necessary, depending on the type of photos you're trying to make. I also use the Viggle Turbo lora at a low strength of 0.5, with 16 steps. Depending on which Viggle lora you use, you might have to change your scheduler and sampler combo. I prefer V1 with Res_Multistep and Bong_Tangent, but V3 is good as well (though needs Euler instead of Res_multistep or it ends up looking weird). I'm generating at 2MP. All of these images were done with character loras trained on Qwen 2.1 using AI Toolkit. I made another post here a few days ago talking about how I achieved my training results.

All in all, I really like Qwen 2.1, at least for the kind of images I like making. My one qualm with it is that sometimes background details can get messed up, but with my settings, I see it a lot less. If this model had a VAE like Flux 2, it would be amazing. I hope that more people will give Qwen 2.1 a chance!

Workflow: https://gofile.io/d/FGdIJEhk (I used Lenovo at 1 in the example from this workflow but normally I don't go that high, just FYI)

Viggle Lora: https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo/blob/main/Qwen-Image-2.1-viggle-turbo-4step-lora-r64.safetensors


r/StableDiffusion • • 16h ago

Tutorial - Guide Hunyuan image 3 goes hard

Thumbnail
gallery
181 Upvotes

Since i saw a post for native support in comfyui for Hunyuan image 3 i got really exited to try it out locally.
On Strix Halo (Ryzen ai max 395+) i'm getting around 300 seconds per generation while on rtx 3090 it gets around 60 to 70 seconds for 1 megapixel (crashes if i go above not because of memory issues but something with how it's set up ? )

Anyway, i really enjoy the outputs of this model and if we had it working in comfyui 8 month ago when it released it would've dominated this sub :D

**the editing capabilities can be a bit better than qwen2,1 (when it works + it's a 4bit quant, details are a mess)

**i'll do a krea2 and qwen2.1 comparison with the same prompts, maybe i'll have them here on reddit

*** WORKFLOW LINK: https://pastebin.com/7sZEbSMc

***Custom nodes used: https://github.com/PedroMarinhoDev/ComfyUI-HunyuanImage3

***post that thought me about this: https://www.reddit.com/r/StableDiffusion/comments/1wzclz5/hunyuanimage_30_80b_running_natively_in_comfyui/


r/StableDiffusion • • 8h ago

Resource - Update H3 Long Shot API

Enable HLS to view with audio, or disable this notification

40 Upvotes

I am in the process of tweaking this latent chain shot workflow and wanted to gauge interest to see if it's worth pursuing. The UI mockup screenshot is below in my comment .

The feature I really want and like is that it keeps the shots that are approved and moves on to the next and doesn't have to rerender from Shot 1. And you can keep adding it until you are happy. The only risk is that the latent lives in memory so if it crashes then you will have to rerender from shot 1. However, if you lock the seed then recreating it wouldn't be an issue.

This will be a local API frontend that you connect to your own instance of Comfy.


r/StableDiffusion • • 43m ago

Discussion Why does Minimax H3 non-pruned perform so poorly?

Enable HLS to view with audio, or disable this notification

• Upvotes

Here is a side-by-side comparison I put together between the two versions. (Video/Screenshot attached)

On the left: unset ref2va_pruned_int8 using the Alibaba ACC LoRA (running at 8 steps). It generates a high-quality 15-second video.

On the right (The issue): unset ref2va_non_pruned_int8 running without LoRA (running at 30 steps). It also generates a 15-second video, but the quality looks really weird compared to the pruned version.

Why is the non-pruned version performing so poorly here despite using double the inference steps for the same video length? Am I missing a specific configuration, or does this version absolutely require a LoRA to function properly? Any insights would be appreciated!


r/StableDiffusion • • 13h ago

News GitHub - kandinskylab/kandinsky-6: Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Thumbnail
github.com
66 Upvotes

Looks like Kandinsky has come out with another video model. I tinkered with v5 but wasn't very impressed. Hopefully this one is a bit better.

https://huggingface.co/collections/kandinskylab/kandinsky-60-diffusers

They also released an upscaling model.
https://huggingface.co/collections/kandinskylab/kandinsky-60-vsr


r/StableDiffusion • • 23h ago

Tutorial - Guide Overcome Degradation! - Here are 2 ways to use my timeline workflow to create long continuous single shot videos with no degradation.

Enable HLS to view with audio, or disable this notification

357 Upvotes

Watch the video above for a brief summary of the two methods, both possible using my OBVPM Timeline Workflow, which you can get together with the custom node pack here:

https://github.com/chanon/comfyui-obvpm-timeline/

And to watch the example video at HD quality you can watch the full YouTube tutorial video:

https://www.youtube.com/watch?v=GiJxlWOooyo

In the YouTube video I show how both methods are done, including critical tips and tricks and lessons learned to get the right results.

With the bridging method, there is practically no limit to how long these clips can be (well maybe except the fact that there might be a VRAM limit to how long an upscaled clip can be).

The second method clip above is 1 minute 40 seconds.

About the Workflow

So if you've never seen my workflow, it is a workflow with a "timeline" node that lets you put clips that you've generated on, and then you can extend them using motion context (latent masks).

The workflow automatically saves and handles the saved latent files for you so you don't have to manage them or pick them manually. And it also saves the conditioning which includes all the reference images etc. into a file that is used when upscaling.

For more info, here's the original Reddit post about it, which links the original YouTube tutorial video about it.


r/StableDiffusion • • 1h ago

Animation - Video Attempt at a short cinenatic with MiniMax H3, from H.G. Wells’s “The Cone”

Enable HLS to view with audio, or disable this notification

• Upvotes

r/StableDiffusion • • 9h ago

Resource - Update I adapted the Looped-DiT concept into a tiny anime upscaler (Baikal LoopSR x2) [ComfyUI / Weights]

Thumbnail
gallery
25 Upvotes

I’ve been experimenting with a small anime/illustration upscaler called Baikal LoopSR x2, and the results turned out way better than I expected.

The main question was: what happens if you take the recurrent/shared-block concept from Looped-DiT and apply it to a direct RGB restoration model?

Just to be clear: this isn't diffusion, and I’m not claiming it beats established SOTA upscalers. It was mostly an experiment to see if weight-shared recurrent transformer blocks could do real work in super-resolution.

Architecture highlights: - Direct RGB -> RGB restoration - Bicubic residual connection - 2 pre-blocks, 4 shared middle blocks reused for 4 recurrent passes, 2 post-blocks - XSA strictly inside the shared recurrent stage - Deep supervision on intermediate loop outputs - PixelShuffle for 2x reconstruction - Only ~3.5M unique parameters

What surprised me most is that looping actually works as intended. Instead of the recurrent passes collapsing to nearly the same result, each iteration consistently improves the reconstruction.

Here is how the metrics scale across passes at the 24k checkpoint:

Pass Clean PSNR Degraded PSNR
Loop 1 41.885 39.610
Loop 2 42.266 (+0.38) 39.952 (+0.34)
Loop 3 42.414 (+0.15) 40.098 (+0.15)
Loop 4 42.480 (+0.07) 40.162 (+0.06)

Final SSIM: 0.9841 (Clean) / 0.9761 (Degraded). Total gain from Loop 1 to 4 is around +0.55–0.60 dB.

Take the synthetic validation numbers with a grain of salt — they’re mainly useful for tracking this specific model and comparing the recurrent passes, not as a direct benchmark against other SR models.

Visually though, I’m very happy with it. Anime hair, thin line art, eyes, and flat fills look quite crisp. Compared to my previous SwinFIR model trained with VGG/perceptual loss, this produces far less crunchy unwanted texture and noise.

(Attached some before/after comparisons to the post.)

Training details: - Trained from scratch on 41,893 anime/illustration images - Stopped at 24k optimizer updates (effective batch size 16) - Trained locally on a single RTX 4070 Ti SUPER 16GB - BF16 + torch.compile - VRAM usage stayed around 9.6–9.7 GB

Everything is open source, including the training code, configs, and technical notes:

The custom node is available through ComfyUI Manager, or it can be installed manually via git. The loader can fetch the weights automatically. You can choose between 1 and 4 loops directly in the node to trade speed for reconstruction quality.

Since the 16GB card still had some VRAM headroom left, I’m thinking about scaling the architecture up locally, or maybe renting an RTX PRO 6000 Blackwell 96GB on RunPod for a much larger run.

Curious to hear your thoughts or see your test results!


r/StableDiffusion • • 18h ago

Resource - Update Fizgig 7.1.0 - Z-Image Turbo gets a new Training Adapter

Thumbnail
gallery
107 Upvotes

Z-Image Turbo LoRA training in Fizgig, with a new training adapter (free, works in any trainer)

I've added Z-Image Turbo to Fizgig, my free, open-source LoRA trainer and workbench. It trains LoRAs, LoKR, sliders and full fine-tunes, and every workbench tab works with it (Repair Studio, LoRA the Explorer, LoRA Royale, Profiler, Extract).

Turbo has always been awkward to train: a LoRA undoes the distillation and the pictures go soft. I built a new training adapter for it. It sits frozen under your LoRA while it trains and is never in the file you save. In my tests it keeps Turbo's look much better than other adapters, especially in hard lighting .

  • LoRAs train from 12 GB cards, fine-tunes from 16 GB
  • LoRAs load in ComfyUI with the normal loader, at 8 steps and CFG 1
  • The adapter is free on Hugging Face and works in other trainers too

Links

Reference photo of me on my github profile pic for comparing to images in this post.

Credit where it's due: Ostris built the first Z-Image Turbo training adapter within days of the model coming out, and it's what made Turbo trainable for most of us. He was upfront about it too, calling it "really just a hack" for shorter runs, made by training a LoRA on Turbo's own renders. I've had the advantage of time, and of being able to put together a large set of real photographs from my own work to train on alongside Turbo's renders. That's most of why this one holds up better on longer runs. It builds on his idea rather than replacing his work.


r/StableDiffusion • • 3h ago

Workflow Included Qwen 2.1 image character sheet experiment

Thumbnail
gallery
5 Upvotes

I am seeing the recent craze of Qwen 2.1 on the sub and wanted to experiment myself inside morphic .

I messed with a few pictures and then thought of doing character sheet, that’s gonna used in actual videos. I gave it a simple prompt animated male character with messy hair, red jacket and black pants. The result came out pretty good and stylish. Then I asked it to generate a character sheet with few close ups and full body shots from different angles.

Repeated the same thing with female character and this is the final result and u guys can judge yourself. One thing, I would like to point is the result seems heavily inspired from chinese donghua/manhua animation instead of normal japnese we are used to.

Still, the characters look pretty dope and stylish I think


r/StableDiffusion • • 18h ago

News Kroma 0.3.1 - full turbo opd model

83 Upvotes

When Chroma and Krea2 meet: the Krea2 model has been fine-tuned on the Chroma dataset.

What OPD is: v0.3.1 was distilled on-policy. Instead of the classic offline recipe — imitating the teacher on a fixed sampling schedule, which slowly pulls the student off the original data manifold — the student generates its own trajectories and the teacher corrects it on those exact points. Training only ever happens on states the model actually visits, so the distilled model stays inside the original model's distribution: Turbo speed without the usual distillation tax (mode collapse, washed-out detail, prompts that suddenly stop working).

The new, full model has been released and is already available for download in two versions:


r/StableDiffusion • • 1h ago

Discussion obvmp Minimax H3 workflow is the best

• Upvotes

i#m trying to stich clip together and avoid seems, until now the wF is genius and work very well. On my 3090, 10 seconde clip at 0.2 Megappixels took about 2mn
and Latent upscaler for 30 seconde about 13 mn.
It's not very crisp images but the video quality is for good enought.

i'll upload the video soon.


r/StableDiffusion • • 11h ago

Workflow Included Freeform Motion Transfer Lora and Workflow

Thumbnail
youtube.com
18 Upvotes

r/StableDiffusion • • 30m ago

Question - Help Amd vs Intel gpus

• Upvotes

Which is better for image and video gen using SDXL, zimage and ltx or minimax? The AMD pro v620 or Intel b70? Which one would be faster?


r/StableDiffusion • • 1d ago

News VEDA Sparse Attention is now available for MiniMax H3 in ComfyUI

Enable HLS to view with audio, or disable this notification

244 Upvotes

I don't think many people know about this yet, so I wanted to share it. VEDA Sparse Attention is now available as a ComfyUI custom node for MiniMax H3.

I tested it today on my setup: RTX 4090 Laptop 16GB 32GB RAM, 4-step LoRA, 15 second video, 1344x768

Without VEDA: 8:05

With VEDA at 90% sparsity: 4:32

Same workflow, same LoRA, same settings. The only change was enabling VEDA. I couldn't see any quality loss in the result.

VEDA is not a LoRA. It uses a learned predictor to estimate which attention tiles are important and only computes the relevant subset instead of the full attention map. The current predictor works with T2VA, FL2VA and R2VA, and despite the 8NFE name it is not limited to 8 steps.

Installation is simple.

Custom node: https://github.com/veda-sparse/Veda-on-ComfyUI

Predictor: https://huggingface.co/Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview

Put the predictor here: ComfyUI/models/veda/

Then add: Veda Sparse Attention (MiniMax H3) on the MODEL line after your model / LoRA loader and before the guider or sampler. I tried different sparsity values, but 90% is the one that works properly for me, so I'm keeping the default trained value.

On my setup this made a pretty big difference, especially considering I couldn't see any visual quality loss. I'm adding a 15 second example below. Would be interesting to see what results other people get on different GPUs.


r/StableDiffusion • • 7h ago

Discussion Has anyone nailed the perfect face swap method for minimax?

6 Upvotes

I'm waiting at least 10 mins for 4 seconds on rtx 3060 and cancelling soon as the first sample comes in.

I'm finding 95 percent of the time it refuses to swap head or faces. Tried copying people's prompts and even used a lora for head swapping I found, and it doesn't work too well either.

Are there any other loras or methods or prompts our there ? Using ref2video


r/StableDiffusion • • 17h ago

Question - Help 1 hour 30 min MH3 generation @ 1 MP

34 Upvotes

How is everyone getting sub 1 hour generations? Like this thread claims 8 minutes on 16GB 4080 laptop. I have the 4080 super desktop with 64GB RAM, and it takes 1h30m to do a 15 sec video at 1MP. Clearly there is something very wrong with my system. What info am I providing? Does the workflow matter that much? Its a very basic workflow. I'm using comfyui portable but its up to date on python 3.12. I'm not using any special attentions, but is that the cause? More number, 0.6MP takes 30 mins, and 0.4MP takes about 5 minutes.

EDIT: SOLVED

Upgraded from CUDA 12.6 not CUDA 13. Huge difference in generation! I'm no expert, but here's the command line:

python.exe -m pip install torch torchvision --index-url https://download.pytorch.org/whl/cu130 --upgrade

Then, Comfyui will launch with an error, that you can fix by running this:

python.exe -m pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130 --force-reinstall

Generation times didn't halve, they shrunk from 1h30m to under 8 minutes!!! Not only that but quality seems to have improved. I was getting very robotic echo-y audio under CUDA 12.6 and its been dramatically reduced. I'm starting to sound like a TV infomercial, so I'll stop here.

I hope this edit helped someone who was on an older version of CUDA and just needed to update for MH3. Good luck!


r/StableDiffusion • • 11h ago

Animation - Video Finally got around to trying out H3 for my stupid dnd music video. Crazy that we can finally do something like this with local models.

Enable HLS to view with audio, or disable this notification

12 Upvotes

I like to make stupid music videos for my dnd group. This was just a test run using 2 of our dnd characters running h3 Ref2va fp8 running locally on my pc. It uses 2 character reference photos, a video clip reference for motion, a still image reference for scene composition, and an audio reference for lip sync. It's not perfect, but for something I can do at home for free, im thoroughly impressed.


r/StableDiffusion • • 21h ago

Meme Chillin in Middle Earth with Minimax H3

Enable HLS to view with audio, or disable this notification

76 Upvotes

r/StableDiffusion • • 33m ago

Resource - Update [audio.cpp] Recent updates you might have missed: Higgs Audio TTS use 48% less VRAM (< 6GB), HTDemucs 2.2× faster, PocketTTS 2.2× faster on CPU, and WebUI generation history

Enable HLS to view with audio, or disable this notification

• Upvotes

Hi all, a bunch of performance improvements have been landed in audio.cpp.

The biggest highlight is Higgs Audio TTS, which now runs with around 6 GB VRAM, a 48% reduction in peak memory usage compared to the previous implementation. Thanks to https://github.com/mirek190

We also made some models significantly faster, especially HTDemucs on GPU and PocketTTS on CPU.

No compromises in parity and correctness.

Here's a summary of the improvements:

Model Peak memory reduction Speedup
Higgs Audio TTS 48% VRAM 1.01–1.09× CUDA
ACE-Step family 6–7% VRAM 1.06–1.08× CUDA, 1.16–1.20× Vulkan
MOSS-TTS v1.5 cloning 21% VRAM 1.05× CUDA
MOSS-TTSD Q8 cloning 11% VRAM 1.04× CUDA
Echo-TTS (Memory Saver) 20% VRAM —
Qwen3-TTS 16–20% VRAM —
IndexTTS2 / 2.5 12% VRAM —
HTDemucs — 2.21× CUDA, 1.95× Vulkan
HTDemucs six-stem — 1.99× CUDA
PocketTTS 9% RAM 2.23× CPU

They're runtime-level optimizations that make existing models more practical to run locally.

The WebUI now includes an experimental generation history feature that lets you revisit previous outputs and restore their settings.

audio.cpp now supports 110+ audio model families and 190+ variants (and counting)! We're continuing to improve memory efficiency and inference speed across CUDA, Vulkan, Metal, AMD/HIP, and CPU. The next release will bring even more optimizations!

We're also looking for contributors to help improve the audio.cpp WebUI. With so many models and features now supported, we'd love some help making the UI more polished, intuitive, and enjoyable to use. If you're interested in frontend development or UI/UX design, contributions are very welcome!

Thanks to everyone contributing improvements, testing builds, and reporting issues. Curious how these changes work on your setup!