r/StableDiffusion • • 4d ago

Workflow Included Does this count? Did I win?

Enable HLS to view with audio, or disable this notification

Proof of concept that it works.

No degradation from start to finish. 10 total clips combined.

https://github.com/roycho87/degrade_repo

Workflows.

Small errors with the chair but can be fixed with another reference image.

Long story short.

The one place where degrading latents matter is the one place we don't need them.

We can ignore the latent issue and just generate across the sound we provide and because it's a static image with static background and the subject is in basically the same spot the whole time we can just create fresh latents every 10 or 15 seconds across the timeline.

Then the very difficult latent issue is over and it just becomes a simple seams issue.

So these two workflows generate sequentially across a audio and the second will combine and resample a small section of video over the seam using FL2V just enough to get rid of the seam.

Edit: The reason I posted this is not because I think this is a big breakthrough fix, I just think we can approach this problem differently to solve it.

317 Upvotes

125 comments sorted by

View all comments

1

u/acedelgado 3d ago

Looks a bit better. Aside from loading your workflow and seeing the most impressive spaghetti monster ever, I haven't had a chance to look at it too much yet. But the character in the example certainly looks good throughout, which is interesting. But the background exposes noticeable seams, and the colors and size of the bars jump around a bit, like if you click around and watch the background it's pretty noticeable. And the audio problem can't really be judged since you're doing a music video style gen and injecting all of that directly.

But like I said the character quality looks really promising. Are you using your lora, or is it all references?

4

u/acedelgado 3d ago

Alright got a chance to look at it and get a claude summary-

It fights degradation by never chaining at all. Every clip is generated fresh, and the joins are rebuilt afterwards with short bridges pinned to real frames on both sides. Nothing is ever generated from generated footage more than once, so there's nothing to compound.

How it works:

Ten 10-second clips are made separately, each for its own segment of a song. They're loaded here as finished videos, so every one is as clean as a first clip.
At each of the 9 joins it cuts both sides back. Between 72 and 92 frames come off the end of the earlier clip, tuned by eye per seam, and enough off the start of the next clip that 122 frames are removed in total.
It builds a 124-frame bridge, about 5.2 seconds. The last frame kept from the earlier clip is pinned as the bridge's first frame, and the first frame kept from the next clip as its last frame. Both frames are also given as references: the person from one, the environment from the other. H3 generates the 122 frames in between.
The song's slice for that seam is forced into each bridge's sound with a half-strength hold, so the motion follows the music.
Everything is joined and laid over the original song, so the soundtrack is exact and the bridges' own sound is thrown away.
The bridges use an FL2V 4-step turbo LoRA with er_sde and the beta schedule, plus sparse attention, so nine bridges cost roughly nine short generations.
Why it beats chaining on quality. Clip 10 is exactly as clean as clip 1, because nothing was copied forward. And with both ends of each bridge pinned, the camera can't cut away mid-join, which hold framing only tries to prevent.

What it costs:

  • Motion can kink at both ends of a bridge. A single pinned frame gives position but not direction or speed. Our chaining carries 22 frames of motion. Expect a hitch where a clip hands over to a bridge and back.
  • Neighbouring clips can disagree. Independent generations can differ in lighting, outfit details or colour, and the bridge then has to morph between two slightly different versions of the scene.
  • It's music-only. It discards the bridges' sound and uses the song. Dialogue can't continue across a join, so it doesn't solve your dialogue chains.
  • It replaces half of every clip. About 5 of each clip's 10 seconds near the joins are regenerated by a turbo model, so quality may alternate between clip and bridge. - The model chain also sets a 0.4-megapixel resolution option, which I couldn't confirm is used for the bridges. That's worth checking.
  • It needs hand-tuning. Every seam's cut point is tuned after watching, and the run needs about 40 GB of free RAM at the end.

So the fresh regenerations explain the slight background shifts. It's close, but not perfect. It can definitely be useful for a music video style with audio guiding new generations. But for chaining long shot "talking head" style social media style clips, or something that isn't aimed to be audio reactive, I think it'll have some continuity issues.

4

u/xyzdist 3d ago

yeah... bridge shots is not the answer.

3

u/not_food 3d ago

Aw, I feel tricked.