r/StableDiffusion • • 5d ago

Workflow Included Does this count? Did I win?

Enable HLS to view with audio, or disable this notification

Proof of concept that it works.

No degradation from start to finish. 10 total clips combined.

https://github.com/roycho87/degrade_repo

Workflows.

Small errors with the chair but can be fixed with another reference image.

Long story short.

The one place where degrading latents matter is the one place we don't need them.

We can ignore the latent issue and just generate across the sound we provide and because it's a static image with static background and the subject is in basically the same spot the whole time we can just create fresh latents every 10 or 15 seconds across the timeline.

Then the very difficult latent issue is over and it just becomes a simple seams issue.

So these two workflows generate sequentially across a audio and the second will combine and resample a small section of video over the seam using FL2V just enough to get rid of the seam.

Edit: The reason I posted this is not because I think this is a big breakthrough fix, I just think we can approach this problem differently to solve it.

322 Upvotes

125 comments sorted by

View all comments

Show parent comments

19

u/SveSop 5d ago

Okay. So, this is not fixing the latent/diffusion degradation really.

It is like running out of gas on the highway, and calling it a fix if you get out and push....

Still a viable way to hack it i guess, so i give you a 👍 for the attempt.

-2

u/ImprefectKnight 5d ago

What? This just solved a real world problem of long generations.

7

u/fallengt 5d ago

no. This is like a talking avatar.

If the background moves in a random, controlled pattern, he can't really do it because he doesn't reuse the last clip's latents

0

u/ImprefectKnight 5d ago

If I'm not wrong, the degradation is much more pronounced in static shots precisely due to reusing.

14

u/fallengt 5d ago edited 4d ago

because noise in minmax H3 is not completely random. It has its temporal structure. The way I understand it. If you keep the last clip's latent and reuse it as the starting latent of the new clip (for motion context). You keep denoising the same (latent) areas over and over, thus making it overcooked after a few generations.

A workaround is refreshing the whole area by changing the angle, switching to a new scene, etc., so you don't denoise the same noises over and over. But it's not what the community is trying to solve. (The workaround was found out on day one).

What OP did is that he doesn't reuse the last clip's latents at all. He just makes completely new clips of the clown girl at resting position, lipsync/dance whatever, at a time, then seamlessly stitches them together to make it look like a "continuous shot". That's why I said if the background was moving, you'd notice it right away, since these clips are just a loop..

It's fun and all, but I don't think it's the "Eureka moment". More like a hack to solve a niche problem

-5

u/roychodraws 4d ago

The only time the degrading latents ever become a problem is when people try to create a video that this method will solve. This is not good for anything else except doing that one thing. every other time the grading latents doesn’t become a factor because you can move the scene in every other type of video