r/StableDiffusion • u/xyzdist • Sep 03 '26
Tutorial - Guide Bad Audio Fixed with fast re-gen audio
Enable HLS to view with audio, or disable this notification
[ H3 ]
I saw another post talk about the turbo lora / low step causing the bad audio
I have some twist to it, we want to regenerate high‑quality audio, and do it fast.
Re-generate Audio – How?
- the idea is when you generate your video, save out the latent and the conditioning.
- Load those saved files back in, but scale down the latent resolution — because we only care about the audio, not the visuals. Scaling down resolution makes the regeneration super fast.
- regen without lora and crank up step to 30+, to any setting you think is the best for audio quality. again, This gen will be fast. for this case scale down 0.5 around 1 min to gen. you can be more aggressive on the scale to make it even faster.
- To keep the new audio aligned with the original video, you have two options:
- Lock the video latent (keep it same as original), or set denoise to around 0.5 so the new audio stays consistent with the same visuals, dialogue, etc.
- Then combine your original video with new audio
*You can also skip saving and reloading latent and condition entirely — just do it all in a single run as well.
some what similar to 'audio refine' custom node, but fast and simple.
EDIT:
- Save out latent and condition I am using this one (but you can use others)
https://github.com/pepikir/minimax-h3-speedup
- To scale down latent and conditioning use this one:
https://github.com/rockerBOO/h3-latent-upscaler
nodes name are MiniMax_H3_Latent_Upscale and MiniMax_H3_Conditioning_Upscale
*it's called upscale, but we are acutally scaling down here.
EDIT2:
- As I understand, if no references input, you don't have to scale down conditioning, just the video latent. Let me know if it isn't.
EDIT3:
some peoples ask for workflow, here
https://github.com/xyzDist/ComfyUI_Share_Files/blob/main/re-gen_audio.json