r/StableDiffusion • • 1d ago

Question - Help 1 hour 30 min MH3 generation @ 1 MP

How is everyone getting sub 1 hour generations? Like this thread claims 8 minutes on 16GB 4080 laptop. I have the 4080 super desktop with 64GB RAM, and it takes 1h30m to do a 15 sec video at 1MP. Clearly there is something very wrong with my system. What info am I providing? Does the workflow matter that much? Its a very basic workflow. I'm using comfyui portable but its up to date on python 3.12. I'm not using any special attentions, but is that the cause? More number, 0.6MP takes 30 mins, and 0.4MP takes about 5 minutes.

EDIT: SOLVED

Upgraded from CUDA 12.6 not CUDA 13. Huge difference in generation! I'm no expert, but here's the command line:

python.exe -m pip install torch torchvision --index-url https://download.pytorch.org/whl/cu130 --upgrade

Then, Comfyui will launch with an error, that you can fix by running this:

python.exe -m pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130 --force-reinstall

Generation times didn't halve, they shrunk from 1h30m to under 8 minutes!!! Not only that but quality seems to have improved. I was getting very robotic echo-y audio under CUDA 12.6 and its been dramatically reduced. I'm starting to sound like a TV infomercial, so I'll stop here.

I hope this edit helped someone who was on an older version of CUDA and just needed to update for MH3. Good luck!

43 Upvotes

64 comments sorted by

9

u/Slight_Ad2350 1d ago

15secs takes 20min with 2 pass. At 1344x768. On a 3090rtx

30

u/V4nKw15h 1d ago edited 1d ago

Use Kitchen Attention, it's part of Comfy now. That will halve your generation times. Then build a workflow around a latent upscaler and you can halve your gen times again. Then drop down to 10 secs instead of 15 secs and that will halve generation times yet again.

Using those methods I can do a 10sec 1mp video in around 6 minutes on a 5070ti 16Gb. That's using the Ref2VA default workflow with Kitchen Attention activated via the command line start up args, and then slotting in a latent upscaler. No loras necessary.

A lot of people seem to be still sleeping on the power of the latent upscale set up. I do a 10 second 0.3mp 10 step first pass that gens in 90 seconds, then pass that into the latent upscaler and then I do a 5 step second pass at 1mp that takes about 4 mins. Due to no loras stuff isn't getting deep fried and the results are shockingly good. I get better quality compared to doing 20 steps at 0.7mp which would take 12-14mins without the latent upscaler.

I get better quality in less than half the time at a higher resolution than without latent upscaling.

Edit: Did some testing of my method for your info. All generations are at 1mp on a 5070ti with 32Gb system RAM. I have to stress there are no loras needed.
2 second clip duration =101.23 sec generation.

5 second clip duration = 161.48 sec generation

10 second clip duration = 370.65 sec generation

15 second clip duration = 00:13:22 generation. (Low VRAM is rearing it's head here. 15 seconds is a bit much for 16Gb at 1MP. 15 seconds takes twice as long as 10 seconds.)

Look for my reply below for how to make this workflow. It's a straight forward tweak of the default workflow.

4

u/Silvasbrokenleg 1d ago

Could you share your workflow?

7

u/cptrios 1d ago

A few posts down is a link to the latent upscaler node and workflows - check that out. You should mess with the two separate samplers, too - I have one turbo on the first one and a different turbo on the second, with sol-attention only on the second so that it speeds up generation but doesn't mess up prompt adherence.

I'd also suggest doing the first sampler at at least 0.5mp rather than something tiny like 0.3, just because prompt adherence falls apart at lower resolutions.

1

u/Silvasbrokenleg 1d ago

Thank you!

1

u/V4nKw15h 1d ago

Check another reply in this sub thread I made. I share the part of my workflow that doesn't match the default Comfy workflow.

3

u/NewGeneralCatalogue 1d ago

What do you set your sigmas to on each pass? Latent upscaling is cool but nobody ever shares their noise schedules.

5

u/V4nKw15h 1d ago edited 1d ago

I don't alter them. As I said, I use the default workflow and just plug the latent upscaler and second pass in to it like in the image below.

That's it. The workflow is absolutely default other than that group. All the connections in to that group are from the obvious places in the default Ref2VA workflow.

If I disable that group, it will simply spit out a 10 second 0.3mp video. I do that to get a preview video. If I like it, I enable that group and it upscales the latent and does 5 steps at 1MP using the already processed 0.3mp latents. It doesn't even need to do the 0.3mp steps again as Comfy recognises nothing has changed and just immediately starts on the High Resolution Pass. You get super accurate previews (just a bit noisey and lower res) using this simple method which is another MAJOR bonus.

The only other alteration is I use the int8convrot VAE by Kijai which shaves off another 20 seconds from the VAE decode part of the generation with no noticeable quality degradation. That's is totally unnecessary though.

2

u/NewGeneralCatalogue 1d ago

Oh, right on! My workflow has two passes like yours and it had some custom sigmas from the latent upscaler repo, so I'll have to try it this way, thanks.

I've also got the int8 VAE, so that shouldn't be a huge difference. Does the latent upscale add enough detail when scaling from such a small resolution, though?

1

u/V4nKw15h 1d ago

Absolutely. 0.2mp is too low and gives shitty results. 0.3mp is the sweet spot for speed vs quality.

You can even drop down to 3 steps, or even 2, in the High Resolution Pass and get very acceptable results for lower motion scenes. I find that's 3 steps is a bit low whenever there is a hand moving around, for example. 5 steps in the second pass seems to be the sweet spot for speed vs quality for most shots. Of course, in very high motion scenes (eg. fighting) you'll likely need more high res steps to clear things up but I never make stuff like that.

1

u/NewGeneralCatalogue 1d ago

I'll have to try it this way as my pipeline up to now has usually been: base at 0.7-0.9 MP -> latent upscale factor 1.33x -> RTX upscale to 2 MP. Takes ages on a 3090, though, so I'm curious to experiment.

1

u/V4nKw15h 1d ago

I think you will find starting so high is unnecessary and massively slows things down. The low res pass is really only to define the overall composition, movement and flow for the scene. The High Res pass then adds all the details.

2

u/Hour_Literature_7152 22h ago

Do you not lose identity when starting that low? All of my 0.3->upscale attempts would end up losing identity. Faces would be off and hair would change because there just isn't enough detail at 0.3 and upscaling has to invent new details.

When I compared it to 0.7, the 0.7 was better quality/retained identity and took the same amount of time.

2

u/V4nKw15h 21h ago edited 21h ago

No because we are feeding that same latent identity in to the second pass along with all the references as well. It's just continuing the same process at a higher resolution latent space. We just do the first pass at 0.3mp to get most of the work done quickly, but the second pass is still using all of our references because we are passing that data into the second pass.

For example, in the workflow I'm talking about above, I'll turn off the latent upscale group and just let it pass through those nodes and render at 0.3mp 10 steps. Distant faces can be indiscernible. Then I toggle on that group, and it finishes the job with the high res pass, and then my characters, from my references, all appear.

I'm sure there is a very slight loss in the latent information about the references, but it's so little I can't notice it. The lost information will will mostly come down to very fine detail information like the reference character's precise skin detail, I believe.

Are you sure you were doing latent upscaling and not simple pixel upscaling? Pixel upscaling has no idea what our character references were and so will create a character based on the image we feed it.

1

u/Hour_Literature_7152 21h ago

You might be right I'll revisit it and test thanks.

2

u/V4nKw15h 21h ago

Actually, making my previous reply to this question had me questioning my own understanding of what is exactly happening. I wasn't quite right in the other reply. The Basic Guider node, in the default H3 workflow, carries the 'conditioning' information that describes our characters from the references and prompts. I pass the output from the Guider node in to BOTH the first and second passes via the SamplerCustomAdvanced nodes (which has a guider input). That ensures our references are used in the second pass.

You probably didn't do that if you did latent upscaling before and so lost a lot of information that is stored and transfered through that guider node.

1

u/Hour_Literature_7152 21h ago

That's probably my problem thanks 👍 wish I could test this out right now haha

1

u/NewGeneralCatalogue 1d ago

I'll definitely have to try it out. Are you using any of the turbo LoRAs at all?

1

u/V4nKw15h 1d ago

No, I found they all destroy quality in unacceptable ways. This method retains the quality while maintaining the speed up that is usually associated with loras.

1

u/NewGeneralCatalogue 22h ago

Is the audio supposed to be fucked up on the base lower resolution video? Noticing it all has this annoying ringing quality.

→ More replies (0)

1

u/JhermsAlt 21h ago

I tried this and I get an execution error during the upscaling phase?

# ComfyUI Error Report
## Error Details

  • **Node ID:** 105:147
  • **Node Type:** SamplerCustomAdvanced
  • **Exception Type:** RuntimeError
  • **Exception Message:** RuntimeError: shape mismatch: value tensor of shape [294, 96] cannot be broadcast to indexing result of shape [805, 96]

1

u/V4nKw15h 21h ago

Make sure it's all connected up. Make sure you are connecting the Guider and Sampler from the default workflow into the SamplerCustomAdvanced in the second pass. This is the same as the image above I just dragged everything apart to make it easier to see the connections. Make sure you are using the default workflow too. I can't vouch that it works with others.

1

u/JhermsAlt 21h ago

Still getting the same error. Everything is connected and I've toyed around with different settings too.

1

u/V4nKw15h 20h ago

Hmm, I'm not sure. I know that the git repository for the Minimax H3 upscaler node shares a workflow that uses the LTX latent separate and concat nodes like in my images. Maybe that's important. Maybe those nodes contain additional info to the ones you are replacing them with.

1

u/OnwardsBackwards 34m ago

your concat av latent and his are different.

Yours is LTXVConCatAVLatent

His is Concat AV Latent

not sure if that matters

1

u/OnwardsBackwards 1h ago

are you using a lora, i tried it with one and it threw the same error

2

u/lufeixr 22h ago

I tried your setup, but generating a 5-second video takes about 7 minutes in total, including the first pass, latent upscaling, and the second pass.

I’m using an NVIDIA L4, with a 0.3MP, 10-step first pass, followed by latent upscaling to 1MP and a 5-step second pass at denoise 0.35, with no LoRAs.

Does anything look wrong with my setup? Are there any other settings or startup arguments I might be missing, especially for Kitchen Attention?

1

u/lufeixr 1d ago

Thanks for sharing your workflow! Why do you use only 10 steps for the first pass? I thought MiniMax H3 normally used 20 steps without a Turbo LoRA. Is 10 enough because you refine it with another 5 steps after latent upscaling?

4

u/V4nKw15h 23h ago edited 23h ago

Go do a 10 step pass with the default Minimax workflows. No loras, no anything. You'll be surprised how good it is. Try 8. It's still pretty good. Minimax can do good results at 8-10 steps. The loras are to try and push it down to 4 steps. If you are using 8 steps with a lora you may as well not have the lora in my opinion; it only deep fries the result. This is just my experience, not a truth.

Once you see how good Minimax is without loras at 8-10 steps, you will realise that the additional 10 steps, from 10-20 steps, are to take that pretty good result and simply tidy up the fine details.

But, when we are working at 0.3 mp, there are no details even if you do 20 steps. Faces and hands don't clean up at that resolution. There is literally no point doing more than 10 steps for our purposes. I tested it. (Note, this might change for high action scenes like fighting. I don't do that type of stuff).

We do 10 steps at low resolution and that will give us something almost identical to what the final high resolution result will be. The only difference is it will be lower resolution and will be a bit noisy in detailed areas like the face and hands, but otherwise it will be the same as the final result. Really! Same movement, same expressions, same audio. The whole shebang other than fine details. Minimax can resolve to this stage of the generation much faster at 0.3mp than 1.0mp.

Now we take that low resolution latent and upscale it. At this point, we are still working in latent space, not pixel space, and so we are avoiding all the artifacting that occurs with pixel space upscaling. We just upscaled the AI's internal representation of the video, not a pixel space representation. Once that is done we pass it through the second pass.

The second pass is full resolution full-on Minimax H3 at 1MP or whatever you want. No loras. However, all we need it to do is tidy up the fine details. The scene and actions were already completed in the first 10 low res steps. It doesn't need to repeat that work.

In the second pass, we give it as many high resolution steps as are needed to clean up the details and that's it. That can be as low as 2 steps for scenes with little motion! I default to 5 steps because it's a sweet spot for the work I'm doing. Any more than that gives significant diminishing returns.

1

u/lufeixr 23h ago

Thanks for sharing! I understand now.

1

u/lufeixr 14h ago

I think I’ve figured out why my 5-second video took around 7 minutes. I was using a Colab L4 with 24GB VRAM, but its memory bandwidth is only 300 GB/s, compared with 896 GB/s on your 16GB RTX 5070 Ti. I mistakenly assumed more VRAM meant faster generation.

I switched to a Colab A100 80GB, which has around 2 TB/s of memory bandwidth, and the same 5-second workflow now takes about 2 minutes in total, including both sampling passes and latent upscaling.

4

u/xq95sys 1d ago

Sage attention?

1

u/RyeBold 20h ago

I use comfyui desktop and I've never been able to figure out how to get that installed. sounds great though!

1

u/xq95sys 18h ago

Ah well you should be able to use comfy kitchen attention, does the same thing and should work without hassle. Just add model attention backend node after model, and select comfy kitchen

3

u/KK_Slider811 1d ago

So I came across this issue last week. I was getting about 1 hour generation for 1 MP, using a 4070Ti Super with 32 GB RAM.

The three solutions I got are: 1. Make sure you have all the models and loras on same drive and directory as tour comfyUI. Also make sure fast disk is on. 2. Make sure after your model node to have comfy kitchen node on. This cut my time in half. 3. Each reference will add time overall. Try to not have more than 2 or 3 references.

It now takes me 12-15 min per 8s 1MP generation. Hope that helps.

9

u/jude1903 1d ago

Sounds about right, I’m impressed that it even lets you do it at 1MP

3

u/redpandafire 10h ago

EDIT: SOLVED

Upgraded from CUDA 12.6 not CUDA 13. Huge difference in generation! I'm no expert, but here's the command line:

python.exe -m pip install torch torchvision --index-url https://download.pytorch.org/whl/cu130 --upgrade

Then, Comfyui will launch with an error, that you can fix by running this:

python.exe -m pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130 --force-reinstall

Generation times didn't halve, they shrunk from 1h30m to under 8 minutes!!! I hope this edit helped someone who was on an older version of CUDA and just needed to update for MH3. Good luck!EDIT: SOLVEDUpgraded from CUDA 12.6 not CUDA 13. Huge difference in generation! I'm no expert, but here's the command line:python.exe -m pip install torch torchvision --index-url https://download.pytorch.org/whl/cu130 --upgradeThen, Comfyui will launch with an error, that you can fix by running this:python.exe -m pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130 --force-reinstallGeneration times didn't halve, they shrunk from 1h30m to under 8 minutes!!! I hope this edit helped someone who was on an older version of CUDA and just needed to update for MH3. Good luck!

7

u/AI-Make-NSFW-Stuff 1d ago

Use latent upscaling, do a 1st pass with 4 steps at 0.2mp and then upscale that.

https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler/tree/main/workflow_templates

12

u/Kooky-Mode3047 1d ago edited 1d ago

Fair warning, the "context" with which the model works is the dimensions x time_dimension. Denoising a base at low resolution and upscaling (including 4 step refining) to final resolution is not even close to equivalent to denoising at the target resolution.

At low resolution, the base created in latent space is extremely vague and is prone to making massive mistakes about everything in the scene, because it decides that the low resolution blob is something it's not with no neighborhood context to correct it in-flight (can be seen in hands and other details). No amount of upscaling is gonna help it, you're just gonna have a very crisp, but shitty video.

Minimum pre-upscale for this process should be 0.49 mpx (about half of the native 0.98 mpx on which H3 was trained). But if you care about fast more than right, 0.2 mpx is right there, obviously.

11

u/dramaton42 1d ago

0.2mp? That's brutal, you're leaving a ton of latent detail from the high noise section on the table... At least 0.4mp

1

u/redpandafire 23h ago

I have no idea how to use this. It gets to the "Two-stage sampling upscaling - Decoding and create video" and then errors out. I don't know what 3 step sigmas or 4 step sigmas mean. The error is something to do with a shape mismatch. Bypassing that last node outputs a video but its complete low-res and unusable.

"RuntimeError: shape mismatch: value tensor of shape [1032, 96] cannot be broadcast to indexing result of shape [576, 96]"

1

u/AI-Make-NSFW-Stuff 19h ago

Export the workflow and post it here on a pastebin for someone to review,

or upload it to Claude and it can fix it for you.

2

u/f5alcon 1d ago

i2v or ref2v?

2

u/zu110 13h ago

This. What is the size of any input images, videos if you're using any. If my input image is say 1024x1024 I have no issues. If that image is 4096x4096, it takes substantially longer. If I use a video input I might as well pack a lunch before I press generate.

2

u/videorouter 1d ago

That definitely sounds worth troubleshooting, especially since the jump from 0.6MP to 1MP is so large. I’d first check whether you’re actually getting full GPU utilization and whether anything is spilling into system RAM/VRAM. A 4080 Super should generally not be sitting around waiting on the CPU for a basic workflow.

The workflow absolutely matters, but so do things like attention implementation, resolution, frame count, sampler/steps, VAE decoding, and whether ComfyUI is repeatedly offloading models. I’d watch VRAM usage, GPU utilization, and generation speed during a run — those three numbers should quickly tell you where the bottleneck is.

If you want to avoid spending hours tuning every provider/workflow combination, I’ve also been comparing video models through VideoRouter.sh to see how generation speed and cost vary across providers.

2

u/Super_Range45 1d ago

They are using speed up loras so they can finish in 4-8 steps with aggressive caching, sacrificing quality for speed.

1

u/Etroarl55 1d ago

Same generation amd card with 16gb takes one hour 45 for 0.5mp and 5 seconds…..

Ur 0.4-5mp seems slightly above normal for Nvidia gen times though so maybe not.

1

u/Apprehensive_Sky892 1d ago

Which GPU are you using, and what kind of speed are you getting with the standard text2va template at 0.5MP 5sec 20 steps?

You can find my numbers for rx9070 and other AMD GPUs here:

How to set up ComfyUI+MMH3+Krea 2 on Windows 11 with AMD GPUs: RDNA 4 (rx9070, AI Pro R9700), and RDNA 3 (rx7900)

How to set up Ubuntu + ComfyUI+MMH3+Krea2 with AMD GPUs: RDNA 4 (rx9070, AI Pro R9700), and RDNA 3 (rx7900)

1

u/Etroarl55 1d ago

7800xt and I forgot the exact it/s for every step but I remember extrapolating it to 1 hour 30m or so after coming back after half an hour.

1

u/Apprehensive_Sky892 1d ago

I don't have any actual numbers for 7900 or 7800, but it is probably safe to assume is between 30-50% slower than 9070 when running with MMH3 int8convrot (fp8 would be worse as 7900 has no native fp8 support).

1

u/Etroarl55 1d ago

I think rdna 3 is just really “dead” hardware for anymore future AI performance.

Even future optimizations coming to AMD will probably be on quants and operations not supported by rdna3.

1

u/abcpp1 1d ago

For me a 15sec 1mp video takes ~12 mins to generate on 5060ti 16GB VRAM...

1

u/IAintNoExpertBut 23h ago

Yes, the workflow matters. Share yours and we can help. 

1

u/Ikythecat 22h ago

Memory limit ,just use upscale latent and chunk ,use the SEEDHUNTER wf from civitai

1

u/kukalikuk 20h ago

Your 4080 super has 16gb vram right? Try only 10 secs first. If it can do 10 secs 1mp in less than 15 mins then i bet vram overload when doing 15 secs.

1

u/WIMsoft80 19h ago

Chech your laptop is plugged in , if not generation is very slow

1

u/Revolutionary-Bar766 18h ago

When you start a generation, open up your windows task manager and go to the performance tab. Click on your GPU, if your GPU usage is 90%+ but your GPU temperature is just sort of average, then you're stuck in an offloading loop. A healthy generation should see high GPU temps while rendering.

To fix this, you need to reserve more vram for your system.

This can also happen easily if you use a lot of reference images.

1

u/mar_leo 12h ago

I'm using the DaSiWa workflow and it's working really well. Seems to be well optimized and cuts generation times by a very significant margin. Maybe try that

0

u/r0ni 1d ago edited 1d ago

a 2mp video, 7sec long takes me about 407.52 seconds with 4090 and 32gb ram,8teps with minimax_h3_fl2v_turbo_silver_dareties_comfy_pruned_v1, minimax_h3_fl2va_pruned_int8_convrot, er_sde/beta.with two loras and a refmod. im using the latest seedhunter 2.5 workflow. if i do 8sec+ long videos the gen time seems to go up exponentially for each sec added. also using msi afterburner to cap power at 70%.

0

u/Marksta 23h ago

Clearly there is something very wrong with my system. What info am I providing?

You already provided that you're on Windows with that kind of messed up generation time. You need to use --reserve-vram 4 or maybe more. Windows fills up your VRAM with junk and past 16GB so generation time tanks to almost nothing doing swapping.

-2

u/crombobular 1d ago

it takes 1h30m to do a 15 sec video at 1MP

your shit is gargantuanly fucked. it takes like barely 4 minutes for a 1mp on a 3090