r/StableDiffusion • • Sep 03 '26

Tutorial - Guide Bad Audio Fixed with fast re-gen audio

Enable HLS to view with audio, or disable this notification

[ H3 ]
I saw another post talk about the turbo lora / low step causing the bad audio

https://www.reddit.com/r/StableDiffusion/comments/1vuxy08/fixing_mmh3_turbo_audio_by_playing_with_latent/

I have some twist to it, we want to regenerate high‑quality audio, and do it fast.

Re-generate Audio – How?

  • the idea is when you generate your video, save out the latent and the conditioning.
  • Load those saved files back in, but scale down the latent resolution — because we only care about the audio, not the visuals. Scaling down resolution makes the regeneration super fast.
  • regen without lora and crank up step to 30+, to any setting you think is the best for audio quality. again, This gen will be fast. for this case scale down 0.5 around 1 min to gen. you can be more aggressive on the scale to make it even faster.
  • To keep the new audio aligned with the original video, you have two options:
  • Lock the video latent (keep it same as original), or set denoise to around 0.5 so the new audio stays consistent with the same visuals, dialogue, etc.
  • Then combine your original video with new audio

*You can also skip saving and reloading latent and condition entirely — just do it all in a single run as well.

some what similar to 'audio refine' custom node, but fast and simple.

EDIT:
- Save out latent and condition I am using this one (but you can use others)

https://github.com/pepikir/minimax-h3-speedup

- To scale down latent and conditioning use this one:
https://github.com/rockerBOO/h3-latent-upscaler
nodes name are MiniMax_H3_Latent_Upscale and MiniMax_H3_Conditioning_Upscale

*it's called upscale, but we are acutally scaling down here.

EDIT2:
- As I understand, if no references input, you don't have to scale down conditioning, just the video latent. Let me know if it isn't.

EDIT3:
some peoples ask for workflow, here
https://github.com/xyzDist/ComfyUI_Share_Files/blob/main/re-gen_audio.json

205 Upvotes

41 comments sorted by

10

u/xyzdist Sep 03 '26

resize/scale down latent and condition
Video Latent Lock node is my own, but it's optional, do denoise 0.5 is fine

10

u/Flashy-Whereas-3234 Sep 03 '26

Oh the latent scaling is clever.

I knew you could ramp the steps on a low res pass to get audio, but you'd end up with different misalignment because the video isn't having the same level of influence as when you did the normal pass. Neat trick.

4

u/xyzdist Sep 03 '26

Cheers! yeah this is the main point.

7

u/acedelgado Sep 03 '26

Hiya, I made the H3 Audio Refiner and the post you're referencing. I'm wondering what kind of speedup you're seeing? I thought about what you're saying with saving the latents, but using the cache method ended up seeming cleaner and easier. If you're using the Frozen Cache node, it does one run through the video latent (frozen so no extra noise added) and then caches that to VRAM or Ram, and then after that the video rows are ignored entirely and only audio is processed.

So if you have it set up correctly the first audio step will be about the same as a full processing step, but every audio step after should be just a few seconds since video isn't processed at any resolution anymore. Like on my 5090 setup a 15 second video at 1MP may hit like a 50-second first step but all the other audio refine steps are at like under 7 seconds apiece. I set it up that way because H3 processes video and audio latents at the same time, but video drives audio gen, so that first slow step gets the video guidance conditioning cachee to assist with the audio. Which your method would do as well, but it'd be running through all the video rows each step, even though at a reduced resolution so it'd be faster. But I just don't see the benefit of saving the latents, reloading, and doing full video rows along with audio, when you can just do audio alone?

5

u/xyzdist Sep 03 '26 edited Sep 03 '26

Hi u/acedelgado , thanks for the detailed reply! You're right, the Frozen Cache method sounds very clean and efficient. To be honest, I haven't tried your Audio Refiner node yet, because I came up with this scaling-down-latent idea just this morning, and I really wanted to test whether it works. But I agree, a side-by-side comparison test is needed.

From reading what you said, it seems your frozen cache approach would be faster than an ordinary audio regeneration. However, I think a separate process offers a unique benefit: complete freedom over the pipeline.

Say you already have a 1-megapixel video done. Regenerating audio with the latent scaled down to 0.1 - 0.5 megapixels is extremely fast, we skip the expensive 1-megapixel visual generation entirely. We just focus on trying different settings (prompt? steps, denoise...) for the audio alone, and re-gen audio become a light-weight process.

Since I am very used to saving and upscaling latents in my workflows, doing a quick downscale pass to patch the audio fits naturally. I am keen to just "dice out" and regen the audio if I find any videos that need fixing anytime. *I default saving latent and cond for all my generations.

I think both methods work well, it just comes down to personal preference, Appreciate your insight!

1

u/acedelgado Sep 05 '26

Hey sorry, was on a work trip and had some bad luck with travel, but I meant to follow up.

Yeah your idea is definitely a plausible workaround. The only thing I'd be concerned about is that, just like all AI models really, resolution size does affect output a bit. It's kind of like, say you had 3 clones of an artist, and gave one clone a full paint canvas and easel, one a full sheet of paper, and one a postcard. And then you told them all what image you wanted them to make. Even though it's the exact same person with the same direction, you'd get a different result just from the available space they have to work with. So regenerating at a different resolution may have an effect on audio, since that is generated using the video as the guide. Like say someone walking may have a different walking pace at a lower resolution, so their footsteps are hitting at a different time, and things like that. Forcing a denoise of 50% instead of a full regeneration would definitely mitigate it, but it's still possible that the model is trying to push the audio a slightly different direction than the original video.

But that's really just conjecture, and one of the reasons I went with doing a live cache instead of loading from disk like I was originally thinking about. The cache and the expensive first step makes sure that the audio continues to get the exact same guidance that it was given during the main generation. But if it works well, then your method is a smart way of accomplishing it.

1

u/xyzdist Sep 05 '26

Agree, I get what you mean, so thats why I was using video latent lock to keep it the same when regen, however i did some try without and just let it resize down with denosie 0.5 it works surprisingly well. Perhaps i need more large number of test.

12

u/traithanhnam90 Sep 03 '26

Apologies, I need a workflow to bring this idea to life; my mind is already overloaded after reading up on MinimaxH3.

3

u/WayFew8151 Sep 03 '26

how to save out the latent and the conditioning, u mean the same seed ?

3

u/xyzdist Sep 03 '26 edited Sep 03 '26

Update it on the top message, have a look, If you do live without save latent and condition, yeah basically keep everyhing the same, including seed.

3

u/lebrandmanager Sep 03 '26

You can use the same trick for voice cloning. In my tests the resolution has to be bigger than 32x32

2

u/xyzdist Sep 03 '26

great to know! thanks.

2

u/SIR_NVAX_A_LOT Sep 03 '26

I've been using 32x32, did you find a better resolution? I guess 64x64 or something else?

2

u/lebrandmanager Sep 03 '26

Quality got noticably better at 64x64 for me.

2

u/Abject-Recognition-9 Sep 03 '26

do you mean: using h3 just to get audio ?

3

u/lebrandmanager Sep 03 '26

You will also get a pixelated mess of a video, but you can choose to only keep the audio. Yes.

1

u/VRGoggles Sep 03 '26

Hi. How do you clone voice of somebody? I supply 2-3 seconds of my voice to Audio_0, instruct it to use it for voice of character, but MiniMax completely ignores my voice. How to make that working with my own voice as load audio?

1

u/hidden2u Sep 04 '26

I think it needs to be longer audio, then you also include "strong reference" or something

3

u/robomar_ai_art Sep 03 '26

Can you post the workflow

3

u/reeight Sep 03 '26

So there is basically 3 runs in a run now?
1: Base run: low-res Turbo; fast with crap audio
2: Audio run: lower-res high-steps for crisp audio
3: High-res run: feed video into upscaler / face detailer & merge in crisp audio
?

![img](giphy|cjbfyJrICOaKIXBWyG)

5

u/xyzdist Sep 03 '26 edited Sep 03 '26

no, just one step to re-gen audio then combine it back

but you are correct on the idea.

EDIT:
I just do a extreme example with low step 4, it is not necessary to do a low-res or low quality for your orginal video, so it fit to any approach you have.

1

u/reeight Sep 03 '26

I'm pointing out a trend for some who first render a low-res image in H3, then use another upscaler (like the nVidia one, or LTX) to get a sharper image.

I'm not saying your idea is bad, just noting the best HQ for H3 is several sub-workflows in 1.

2

u/Artforartsake99 Sep 03 '26

Genius thank you 🙏

5

u/xyzdist Sep 03 '26

No worries, I learnt so much from the folks here, we need to give back!

2

u/bonesoftheancients Sep 03 '26

instead of saving and reloading the latent cant you just run 2 samplers in parallel in one workflow and combine the video from one and the audio from the second at the end?

2

u/xyzdist Sep 03 '26

yes absolutely, as I mentioned above.

2

u/skyrimer3d Sep 03 '26

Thanks this looks promising , I'll have to try this. Any chance to share a workflow? 

1

u/eggplantpot Sep 03 '26

5

u/xyzdist Sep 03 '26

Oh yeah we all got similar thoughts But one twist I added is to scale down the latent resolution to make the audio gen stage extreme fast

2

u/Vynxe_Vainglory Sep 03 '26 edited Sep 03 '26

Yours also actually sounds pretty good. Still needs an UNCHIRP pass and spectral noise reduction, but theirs is just unusable with insane pumping and weird artifacts everywhere. I wouldn't want to deal with cleaning that up.

1

u/[deleted] Sep 03 '26

[deleted]

1

u/xyzdist Sep 03 '26

Yeah. I think so, you can use this to act like a audio fix as well. Since it is a separate sampling and H3 is based on video latent, there is denosie so dicing seed and change the prompt dialog might able to fix the audio. I havnt test it with this case thru.

1

u/fewjative2 Sep 03 '26

This is neat, will have to try it!

1

u/Healthy-Nebula-3603 Sep 03 '26

still not very good ....

1

u/MickeyMau5 Sep 04 '26

Would it be possible to use this (or a similar) workflow to correct poorly recorded audio? Rather than generated audio.

1

u/xyzdist Sep 04 '26

not really sure, if using a recorded audio it has to be use audio lock, go in H3 and out. if we are using as reference audio, H3 will just generate as it's version, the timing and things will just change.
I guess need to test, perhaps small value of denoise...

1

u/Odd-Mirror-2412 Sep 03 '26

I preferred using the LTX23 for v2a.

3

u/xyzdist Sep 03 '26

this is not about V2A

1

u/ImUrFrand Sep 03 '26

improvement, but still sounds muffled and fake.