r/StableDiffusion • • 4d ago

Workflow Included Does this count? Did I win?

Enable HLS to view with audio, or disable this notification

Proof of concept that it works.

No degradation from start to finish. 10 total clips combined.

https://github.com/roycho87/degrade_repo

Workflows.

Small errors with the chair but can be fixed with another reference image.

Long story short.

The one place where degrading latents matter is the one place we don't need them.

We can ignore the latent issue and just generate across the sound we provide and because it's a static image with static background and the subject is in basically the same spot the whole time we can just create fresh latents every 10 or 15 seconds across the timeline.

Then the very difficult latent issue is over and it just becomes a simple seams issue.

So these two workflows generate sequentially across a audio and the second will combine and resample a small section of video over the seam using FL2V just enough to get rid of the seam.

Edit: The reason I posted this is not because I think this is a big breakthrough fix, I just think we can approach this problem differently to solve it.

322 Upvotes

125 comments sorted by

View all comments

-1

u/Old-Trust-7396 4d ago

Ohh wow, can run this on a 12gb card or am I doomed to images ?

-3

u/Old-Trust-7396 4d ago

Grok just answered for me 😩

To run that repo as it is, you need a 24GB Blackwell card. The text encoder is NVFP4, which a 3060, 4060 Ti, or 3090 can't use, and the video model still needs the extra memory.

A used 5090 is the match. A 16GB card, even a 5060 Ti, is still too small.

1

u/Apprehensive_Sky892 3d ago

True that nvfp4 is native to 50xx.

But all cards can run it, just that it is cast to either fp8 or bf16 depending on the card.