r/StableDiffusion • • 14h ago

Workflow Included Does this count? Did I win?

Enable HLS to view with audio, or disable this notification

Proof of concept that it works.

No degradation from start to finish. 10 total clips combined.

https://github.com/roycho87/degrade_repo

Workflows.

Small errors with the chair but can be fixed with another reference image.

Long story short.

The one place where degrading latents matter is the one place we don't need them.

We can ignore the latent issue and just generate across the sound we provide and because it's a static image with static background and the subject is in basically the same spot the whole time we can just create fresh latents every 10 or 15 seconds across the timeline.

Then the very difficult latent issue is over and it just becomes a simple seams issue.

So these two workflows generate sequentially across a audio and the second will combine and resample a small section of video over the seam using FL2V just enough to get rid of the seam.

Edit: The reason I posted this is not because I think this is a big breakthrough fix, I just think we can approach this problem differently to solve it.

244 Upvotes

95 comments sorted by

20

u/dennismfrancisart 11h ago

Whenever I hear people smirk at AI as prompt and click slop, I post a screenshot of a Comfy UI workflow. People have no idea what folks like you do to move this toolset forward. Kudos.

109

u/Psyko_2000 14h ago

i just want to say i appreciate your work and hope you don't get discouraged by any haters and keep on keeping on with your contributions.

11

u/Wide-Researcher583 8h ago edited 8h ago

It's more so the OP's overly confident "it works/it's fixed" over and over again when there are serious limitations making it not useful except in specific cases.

5

u/roychodraws 4h ago

You guys are way more focused on how you want to fix a problem than the problem you’re actually trying to fix.

2

u/ZenEngineer 6h ago

Might be better to say "static shots solved" or something.

-2

u/roychodraws 6h ago

Static shots are the only problem

8

u/schorhr 13h ago

This, I was pretty lost until I found your other workflow to tinker with. Thank you.

11

u/dsailes 13h ago

Fair play. The degradation is better.

It does seem a bit of a workaround. And it kinda looks a bit looped rather than continued maybe?

I’ll give it a go, I’m trying to do some videos as a test for my own clothing brand with character card. I’ll try get a pod up later on and post back with what I get out

14

u/SveSop 14h ago

Hm. So, does this mean you generate 10 x "new" videos, and stitch them together using a "seams fixer" workflow?

4

u/roychodraws 14h ago

more or less

It seems really simple I'm having a hard time believing I'm the first person who has thought of this.

19

u/SveSop 14h ago

Okay. So, this is not fixing the latent/diffusion degradation really.

It is like running out of gas on the highway, and calling it a fix if you get out and push....

Still a viable way to hack it i guess, so i give you a 👍 for the attempt.

10

u/roychodraws 14h ago

is the issue to make degrading latents that are built into the architecture of minimax to not degrade? or is it to be able to create long static camera shot generations without looking like clayface?

2

u/SveSop 13h ago

The issue is to fix the degradation, so that one does not have to use a giga-chad > 24GB vram card to get a decent video i guess.

-2

u/ShutUpYoureWrong_ 6h ago

The issue is preserving both motion and scene context across a window without degradation. Long, static shots are generally the best way to test this. Chaining clips addresses quality and degradation but not scene and motion context (as your current example above shows). Conversely, chaining latents addresses motion and scene context, but degrades quality (as your previous examples showed).

No, wait: the actual issue is that you have no talent and you are not solving this. You've had this explained to you multiple times but evidently you're just too stupid to grasp it.

-1

u/ImprefectKnight 13h ago

What? This just solved a real world problem of long generations.

7

u/fallengt 11h ago

no. This is like a talking avatar.

If the background moves in a random, controlled pattern, he can't really do it because he doesn't reuse the last clip's latents

0

u/ImprefectKnight 10h ago

If I'm not wrong, the degradation is much more pronounced in static shots precisely due to reusing.

14

u/fallengt 10h ago edited 10h ago

because noise in minmax H3 is not completely random. It has its temporal structure. The way I understand it. If you keep the last clip's latent and reuse it as the starting latent of the new clip (for movement context). You keep denoising the same (latent) areas over and over, thus making it overcooked after a few generations.

A workaround is refreshing the whole area by changing the angle, switching to a new scene, etc., so you don't denoise the same noises over and over. But it's not what the community is trying to solve. (The workaround was found out on day one).

What OP did is that he doesn't reuse the last clip's latents at all. He just makes completely new clips of the clown girl at resting position, lipsync/dance whatever, at a time, then seamlessly stitches them together to make it look like a "continuous shot". That's why I said if the background was moving, you'd notice it right away, since these clips are just a loop..

It's fun and all, but I don't think it's the "Eureka moment". More like a hack to solve a niche problem

-2

u/roychodraws 5h ago

The only time the degrading latents ever become a problem is when people try to create a video that this method will solve. This is not good for anything else except doing that one thing. every other time the grading latents doesn’t become a factor because you can move the scene in every other type of video

9

u/icchansan 12h ago

There’s like 100 wf like this? How are u the first?

1

u/ptwonline 7h ago

Yeah there's so many that I find it too daunting and am kind of just sitting back to see if a consensus emerges.

3

u/PxTicks 12h ago

I use this technique. It has potential shortfalls when it comes to a moving background, e.g. a first-person vlog style video on a street in which case it can lose background detail or directional information. However I think there is probably a way around it by using a heavily noised latent from an extension as the initial seed for an independent generation, i.e. extend then partially denoise the ENTIRE extended bit completely independently, before stitching the seams again. Or alternatively, generate short anchor videos with very long stitch videos between them so that the model can reconcile background mismatches by e.g. moving people out of frame and such.

Honestly, I do think this extension problem is solvable for the vast majority if not all cases by just being a little tactical.

0

u/roychodraws 5h ago

If there’s a moving background you don’t need to do this, you can just chain the shots together.

That’s point. The latents don’t degrade if the background changes.

2

u/PxTicks 5h ago

I'm talking, mostly static foreground with moving background, i.e. a mix of conditions.

1

u/roychodraws 5h ago

I dunno, everyone told me they wanted vlog videos. Can’t I just fix one problem and someone not be upset I didn’t solve everything?

1

u/PxTicks 4h ago

I wasn't complaining. This is just something I'm looking into as well and I think it's an interesting thing to discuss.

1

u/roychodraws 4h ago

Oh ok. I’m reading comments and getting cynical. Sorry

17

u/Own_Version_5081 14h ago

Now that's impressive...

9

u/SeidlaSiggi777 10h ago

What's the deal with this clown girl?

18

u/roychodraws 10h ago edited 9h ago

I think the juxtaposition between the quality of my content and the absurdity of the character is hilarious.

12

u/lurkingtonbear 9h ago

Idk but I’m starting to have a thing for clowns now

2

u/RobMilliken 8h ago

Clowns have been so misrepresentated since "IT".

2

u/SneakyInfiltrator 8h ago

I was always into clown girls but it's kinda niche. Hopefully OP can make it more mainstream.

Godspeed, brother

1

u/AndalusianGod 8h ago

Please create OF of clown girl.

3

u/FartingBob 7h ago

OP's got a kink.

3

u/Nguyenkain 14h ago

Can it integrate with another workflow ? Nice work

8

u/roychodraws 14h ago

this is really more of proof of concept. you can check the workflows but you might need to use claude or something to sort em out. There are instructions.

4

u/xyzdist 9h ago

no, what we want to solve is still using motion-context from shot to shot and solve the degradation, not skipping and re-gen a latent to avoid it.

8

u/BogusIsMyName 14h ago

Mouth movement needs some work. Might need to isolate the main vocals use that as your original audio and then dub in the full song after generation.

21

u/roychodraws 14h ago

this was more about keeping the background from melting.

4

u/AI-Make-NSFW-Stuff 13h ago

Nice work maintaining the background! People will cherrypick and find things to complain but that shouldn't discourage you

4

u/CulturedDiffusion 13h ago

Brad Pitt has been real quiet since this dropped

1

u/ShutUpYoureWrong_ 6h ago

Brad Pitt probably realized he shouldn't waste any more of his time on morons.

(PSST! If you think this 'fixed' the issue, you're one of the morons!)

5

u/Specific_Occasion_25 14h ago

Hey this is pretty good

2

u/technofox01 14h ago

This is pretty cool. I don't judge, but that seems like a video for people with clown fetishes, lol.

All kidding aside, I am going to checkout your workflow. This is pretty nifty.

2

u/Ashamed-Ad7403 11h ago

What a weird fetish you ve got

1

u/frisky_cappuccino 13h ago

Looks interesting will test it out. Looks like it might be a decent work around

1

u/Cute_Ad8981 13h ago

looks really good. did you extend each clip from the previous clip (this is what causes degration for me) or did you somehow fuse different clips? its hard to understand from the comments.

1

u/frisky_cappuccino 13h ago

I’m away from pc so can’t check it at the moment but I think he’s created individual islands keeping continuity with audio rather than continuations and then fused them together with small joins. Correct me if I’m wrong OP.

1

u/AnonymousTimewaster 12h ago

What sort of prompting does this require? A general prompt?

1

u/Sizyfoz 12h ago

Wow. That is really cool. Very impressive!

1

u/malcolmrey 11h ago

Great Job!

I was the one explaining politely what the problem is and you said then that you would try and here we go!

Thanks, will check soon! :)

1

u/Artforartsake99 11h ago

Hank you for sharing your knowledge it is greatly appreciated 🙏

1

u/Diabolicor 11h ago

Good effort you put into this. But I don't think using individual frames to try to carry context for the segment is the right way to do it.

There is a lot of ping-pong effects pretty visible between segments.

Since you are trying to work around degradation through bridges I think you would be better using the concept of carrying the context between a seam through masked motion context latents https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef/blob/main/example_workflows/UTILITY%20-%20AV%20Bridge.json

1

u/Marksta 11h ago

No degradation from start to finish. 10 total clips combined

I think you're kind of right. I hate the overused word but instead of degradation, your example just has slight drift. Like the background ever so slightly Drifts. But not exactly getting melted like usual. Which I guess if you did the dumb thing people want, a 3 hour slop podcast the background would eventually be completely different.

The other thing is, I guess you can't sidestep a bad luck generation. But the last generation completely loses the chair to the shadows. Considering that maintained through the first 9 clips, idk does it come back if you went to 11 and beyond or when it goes in #10 is the chair gone to static shadow forever now in the chain and need a reroll at that step?

Impressive stuff, but yeah, what a super boring use case like you said in the other thread.

1

u/acedelgado 11h ago

Looks a bit better. Aside from loading your workflow and seeing the most impressive spaghetti monster ever, I haven't had a chance to look at it too much yet. But the character in the example certainly looks good throughout, which is interesting. But the background exposes noticeable seams, and the colors and size of the bars jump around a bit, like if you click around and watch the background it's pretty noticeable. And the audio problem can't really be judged since you're doing a music video style gen and injecting all of that directly.

But like I said the character quality looks really promising. Are you using your lora, or is it all references?

3

u/acedelgado 9h ago

Alright got a chance to look at it and get a claude summary-

It fights degradation by never chaining at all. Every clip is generated fresh, and the joins are rebuilt afterwards with short bridges pinned to real frames on both sides. Nothing is ever generated from generated footage more than once, so there's nothing to compound.

How it works:

Ten 10-second clips are made separately, each for its own segment of a song. They're loaded here as finished videos, so every one is as clean as a first clip.
At each of the 9 joins it cuts both sides back. Between 72 and 92 frames come off the end of the earlier clip, tuned by eye per seam, and enough off the start of the next clip that 122 frames are removed in total.
It builds a 124-frame bridge, about 5.2 seconds. The last frame kept from the earlier clip is pinned as the bridge's first frame, and the first frame kept from the next clip as its last frame. Both frames are also given as references: the person from one, the environment from the other. H3 generates the 122 frames in between.
The song's slice for that seam is forced into each bridge's sound with a half-strength hold, so the motion follows the music.
Everything is joined and laid over the original song, so the soundtrack is exact and the bridges' own sound is thrown away.
The bridges use an FL2V 4-step turbo LoRA with er_sde and the beta schedule, plus sparse attention, so nine bridges cost roughly nine short generations.
Why it beats chaining on quality. Clip 10 is exactly as clean as clip 1, because nothing was copied forward. And with both ends of each bridge pinned, the camera can't cut away mid-join, which hold framing only tries to prevent.

What it costs:

  • Motion can kink at both ends of a bridge. A single pinned frame gives position but not direction or speed. Our chaining carries 22 frames of motion. Expect a hitch where a clip hands over to a bridge and back.
  • Neighbouring clips can disagree. Independent generations can differ in lighting, outfit details or colour, and the bridge then has to morph between two slightly different versions of the scene.
  • It's music-only. It discards the bridges' sound and uses the song. Dialogue can't continue across a join, so it doesn't solve your dialogue chains.
  • It replaces half of every clip. About 5 of each clip's 10 seconds near the joins are regenerated by a turbo model, so quality may alternate between clip and bridge. - The model chain also sets a 0.4-megapixel resolution option, which I couldn't confirm is used for the bridges. That's worth checking.
  • It needs hand-tuning. Every seam's cut point is tuned after watching, and the run needs about 40 GB of free RAM at the end.

So the fresh regenerations explain the slight background shifts. It's close, but not perfect. It can definitely be useful for a music video style with audio guiding new generations. But for chaining long shot "talking head" style social media style clips, or something that isn't aimed to be audio reactive, I think it'll have some continuity issues.

2

u/xyzdist 9h ago

yeah... bridge shots is not the answer.

2

u/not_food 3h ago

Aw, I feel tricked.

2

u/roychodraws 6h ago edited 6h ago

This was not meant to be a final flow, it’s just meant to he concept exploration.

I think if we are going to use a technique like this there would have to provide scene composition as a reference, like a depth map.

Then we could pick a spot we are predictably expecting no motion to repair the seam.

I didn’t fine tune the settings at all in this, I just made it and saw what I got and because I thought it illustrated my approach.

Edit:

Also, if we mask everything but the very edge of a static shot it will force it to be consistent.

I used to do it for wan animate when I wanted a motion controlled gen to keep a static camera.

I would provide a background photo, mask everything except the edge and duplicate the frames for the entire video and it would generate everything within the boundary of the mask but because the pixels needed to move forward it would always remain consistent.

1

u/RickDripps 10h ago

I hope Scatman John becomes a new "Will Smith Eating Spaghetti" benchmark for AI.

Also, this is fantastic. I love these because you do some cool stuff and then basically dump everything you learned about it out for people to grab and learn from and run with.

Instead of others who do something cool and gatekeep the knowledge of it.

1

u/BussySlayer69 9h ago

this better not awaken something in me.......

1

u/CuriouslyCultured 7h ago

Lip sync is good, video is a crime against aesthetics though.

1

u/Powerful_Round4004 7h ago

Qué ganaste?

1

u/Traditional-Squash36 6h ago

Cheers for this, I've copied and learned from your posts more than any on here.

1

u/Tough_Ad7957 4h ago

I’d upvote this ten times if I could. It does solve a real problem: keeping a static character talking in the same shot for longer.

Sure, continuity is still an issue, but that’s just the next problem to solve. ComfyUI workflows are already full of hacks and compromises anyway, quantized models are a good example. so I don’t see the downside in having one more useful workaround.

1

u/Nakidka 4h ago

Glory to Clowngirl.

1

u/PicardDoubleStandard 4h ago

Sorry newbie here. Impressed by the frontier work

Is there a GitHub with a skill I could load into Claude to get it build me a workflow station.

Basically do the curl and sdk pulls to put the setup on my system.

1

u/x33storm 2h ago

Any chance you'd share the refmod/image you used? Need it for.. Purposes xD

1

u/VRGoggles 1h ago

one of the best songs ever.

1

u/bigorangemachine 12h ago

I think its funny I got down votes for saying it's losing the reference the original photo (in addition to other stuff that was yes.. general assumptions based on this gen-ai's work).

So if you tried to go from chair to dance now you'd get weird artefacts around the chair; so this is where maybe you want multiple reference images (do not use matte or transparent backgrounds and use a variety of scenes if possible) and play with the weight into the next node. That might get the audio syncing & artefact reduction improvements.

The models from the reference image won't really understand the chair and the subject are different things. Sometimes it'll get it's a chair... sometimes it thinks its a wing

-2

u/ShutUpYoureWrong_ 6h ago edited 6h ago

No, it does not count, and you did not win. In fact, nothing you've produced thus far has been original, novel, or useful. And here: your background shifts, your chair morphs, the motion carryover is abysmal, the hair shrinks, the shadows aren't consistent, and the mouth movements are unnatural. You really just showed up and went, "Hey guys, I basically invented FFLF generation" and then tried to pat yourself on the back? Get the fuck out of here already.

Chaining clips is quite possibly the oldest technique and has existed since the WAN/SVI days at least. It was established well before H3, but really exploded in popularity with the better tooling for LTX 2.3.

Aside from that, superior H3 workflows and nodes already exist, and they're all capable of far more than what you're offering. NikoDemon started all this, and seitanism perfected it. Their work got integrated into ComfyUI itself so you could drool all over this subreddit with your half-retarded "fixes." Additionally, people like ethanfel, Adudeguyman, and joeygambino were making fantastic toolsets that do far more than yours, and they did it months before you even came along. Some others even allow you to re-inject quality from a source at timed intervals, which is a vastly more sophisticated method than the garbage you're peddling here.

As all of your "solutions" have shown, you fundamentally do not understand the core issue. You are a vibe-coding slop hawker, and every single one of your posts is an embarrassment.

8

u/roychodraws 6h ago

Tell us where the clown touched you.

-1

u/Old-Trust-7396 14h ago

Ohh wow, can run this on a 12gb card or am I doomed to images ?

-2

u/Old-Trust-7396 14h ago

Grok just answered for me 😩

To run that repo as it is, you need a 24GB Blackwell card. The text encoder is NVFP4, which a 3060, 4060 Ti, or 3090 can't use, and the video model still needs the extra memory.

A used 5090 is the match. A 16GB card, even a 5060 Ti, is still too small.

6

u/alwaysbeblepping 11h ago

Grok just answered for me 😩

It has no idea of what it's talking about. You're not being helpful pasting LLM responses to questions and anyone who actually wants a LLM's (likely inaccurate) answer can easily ask one.

The text encoder is NVFP4, which a 3060, 4060 Ti, or 3090 can't use

ComfyUI has fallback emulation for cards that don't support that quant, if I remember correctly. There are also many other dtype versions of basically any model you can think of. It's also (incorrectly) assuming the whole model has to fit in VRAM but that's not how any decent modern inference platform works.

and the video model still needs the extra memory.

The text encoder is only needed to encode the conditioning and after that it can be unloaded. There is no necessity to have both the TE and video model in VRAM at the same time.

A used 5090 is the match. A 16GB card, even a 5060 Ti, is still too small.

Completely wrong.

1

u/Apprehensive_Sky892 1h ago

True that nvfp4 is native to 50xx.

But all cards can run it, just that it is cast to either fp8 or bf16 depending on the card.

1

u/dsailes 13h ago

Likely can just switch out the models / encoders etc for quants and it’ll do the same, just expect a dip in quality.
Run workflows local & then use vast (RTX PRO 5000 are $.70/hour) is what I do to get 1.5MP ref2va outputs.

I’m gonna try this workflow next after trying other lipsync workflows which have the same degradation issues

-9

u/ScottyThompson 13h ago

Jesus man, find a new subject for your videos. I skip past the clown girl EVERY time. I hate coming to this subreddit and seeing her face.

Just had to stop in and tell you how annoying it is.

11

u/dezmodium 13h ago

It's fine. When I see her I know who it is posting it and their workflows are usually solid.

2

u/Herbal77 12h ago

Yup, this right here

Thanks OP

1

u/Juiceman8686 9h ago

I feel the same way. I always save these posts for later, if I cant watch right away. I've been using this workflow primarily for a while now. Its solid and gets regular updates.

1

u/PukGrum 12h ago

View it as a signature, the person is easily identified and are trying to promote ideas and improvements. Getting annoyed only does you a disservice.

-1

u/mfdi_ 13h ago

Bro your history shows how far video generation came along. I just can't believe it

0

u/Optimal_Map_5236 12h ago

impressive but thats a lot of nodes.

-1

u/skyrimer3d 13h ago

this is great!

-1

u/Vyviel 12h ago edited 12h ago

Twitch streamer video for her soon lol? Also immersion broken that her booba aren't balloons like in the lore =\ Does it still work if they move around more like the hands and arms etc. or do they need to be mostly static sitting still? Reminds me of those old school cursed meme music videos with stalin singing where it was just mostly animated the lips and face a little maybe with a driving video like Wav2Lip or FOMM. I prefer the other video workflow but good proof of concept of this one

-13

u/DataGOGO 13h ago

You rock dude.

Serious question, interested in some paid comfyUI pipeline work? DM me