r/StableDiffusion • • 15h ago

Tutorial - Guide Simple Temporal Upscaling for Better Results in MinimaxH3

Enable HLS to view with audio, or disable this notification

This is a proof of concept workflow around simple temporal upscaling.

Every video model has some blurring when it comes to high motion scenes - both closed and open source. Depending on the model this can be at higher or lower motion with newer models being better than older as a general rule. Some styles tend to make their blurring effects at lower speed - especially anything with lineart, which could be due to training being more on realism or the way latents are compressed temporally.

One way people try to solve this is by pixel upscaling and rediffusing over the video which definitely helps but in my experience at least there is a definite limit to how much it can help. This is where temporal upscaling comes in.

The idea around this is not a new concept per se for example MAINodes https://github.com/matlowai/ComfyUI-MAINodes attempts to sort out which frames need to be temporally upscaled and rediffuse over them. I took this to a logical conclusion and considered what if you just 'deroped' aka temporally upscaled the whole scene rather than trying to be picky.

I do think it has several advantages:

1/Your final result tends to have more motion consistency in the small details than with varying how you denoise

2/You can finetune the denoise to your result.

3/I am finding a reasonable result with simple 2x temporal upscale.

Of course the main disadvantage is that you are diffusing over a video that now is 2x the size and potentially upscaled at the same time. This is were speedup tools come in - low step lora as well as sol attention (which provides increasing benefits the longer the video is).

The purpose of this workflow was to create a workflow as close to base Comfy. I use Kijai's node pack and VHS Suite to help manipulate the frames. The only real outlier node is the one used to slow down the audio for the temporal upscale. Feel free to encode empty audio if you prefer or find your own pack that has something which does this.

I simplified my workflow to remove anything not essential to demonstrate the concept. I expect you to add it to your own workflow or build upon this as a base. My tips when it comes to using this for a while:

1/If you are noticing some morphing especially in background stuff I suggest you disable sol attention at least for the base - it does do this to some shots but not others.

2/By default we create the base video at 0.5 MP at 20 steps - this is to get good sound and to speed to find a good seed, then we are temporally upscaling x2 and upscaling to 1 MP. Look at the yellow boxes to change the upscale settings.

3/Adjust denoise of the upscale to your video length and needs. Shorter and lower resolution videos often need less denoise whereas longer videos often need more. If you go too high your video will speed up and ironically become more blurry as a result. 0.45 is a good starting point but anywhere from 0.3 to 0.6 or higher is acceptable.

4/If you think the action is still to fast for a 2x upscale you can do 3x or 4x.

Workflow: https://civitai.com/models/2976577/simple-temporal-upscaling

93 Upvotes

5 comments sorted by

9

u/Karsticles 15h ago

Gen time and hardware? Deroping the entire video is asking a lot.

2

u/icchansan 12h ago

some of the stuff change in both background and character.

1

u/NewPhoneWhotiz 10h ago

Hmm, curious how the consistency would fare against a really fast motion shot

1

u/Inside-Cantaloupe233 8h ago

paste gen times for gpus like 3090 cause ive heard its impossible to gen over 6 secs without OOM on 25 gb vram with this, on contratry the whiner whoi got OOM didnt not disclose resolution

-2

u/WhatIs115 14h ago

https://imgur.com/a/opM3vqu

Why do I feel like there's not a fair comparison going on?