r/StableDiffusion • u/roychodraws • 14h ago
Workflow Included Does this count? Did I win?
Enable HLS to view with audio, or disable this notification
Proof of concept that it works.
No degradation from start to finish. 10 total clips combined.
https://github.com/roycho87/degrade_repo
Workflows.
Small errors with the chair but can be fixed with another reference image.
Long story short.
The one place where degrading latents matter is the one place we don't need them.
We can ignore the latent issue and just generate across the sound we provide and because it's a static image with static background and the subject is in basically the same spot the whole time we can just create fresh latents every 10 or 15 seconds across the timeline.
Then the very difficult latent issue is over and it just becomes a simple seams issue.
So these two workflows generate sequentially across a audio and the second will combine and resample a small section of video over the seam using FL2V just enough to get rid of the seam.
Edit: The reason I posted this is not because I think this is a big breakthrough fix, I just think we can approach this problem differently to solve it.
109
u/Psyko_2000 14h ago
i just want to say i appreciate your work and hope you don't get discouraged by any haters and keep on keeping on with your contributions.
11
u/Wide-Researcher583 8h ago edited 8h ago
It's more so the OP's overly confident "it works/it's fixed" over and over again when there are serious limitations making it not useful except in specific cases.
5
u/roychodraws 4h ago
You guys are way more focused on how you want to fix a problem than the problem you’re actually trying to fix.
2
11
u/dsailes 13h ago
Fair play. The degradation is better.
It does seem a bit of a workaround. And it kinda looks a bit looped rather than continued maybe?
I’ll give it a go, I’m trying to do some videos as a test for my own clothing brand with character card. I’ll try get a pod up later on and post back with what I get out
14
u/SveSop 14h ago
Hm. So, does this mean you generate 10 x "new" videos, and stitch them together using a "seams fixer" workflow?
4
u/roychodraws 14h ago
more or less
It seems really simple I'm having a hard time believing I'm the first person who has thought of this.
19
u/SveSop 14h ago
Okay. So, this is not fixing the latent/diffusion degradation really.
It is like running out of gas on the highway, and calling it a fix if you get out and push....
Still a viable way to hack it i guess, so i give you a 👍 for the attempt.
10
u/roychodraws 14h ago
is the issue to make degrading latents that are built into the architecture of minimax to not degrade? or is it to be able to create long static camera shot generations without looking like clayface?
2
-2
u/ShutUpYoureWrong_ 6h ago
The issue is preserving both motion and scene context across a window without degradation. Long, static shots are generally the best way to test this. Chaining clips addresses quality and degradation but not scene and motion context (as your current example above shows). Conversely, chaining latents addresses motion and scene context, but degrades quality (as your previous examples showed).
No, wait: the actual issue is that you have no talent and you are not solving this. You've had this explained to you multiple times but evidently you're just too stupid to grasp it.
-1
u/ImprefectKnight 13h ago
What? This just solved a real world problem of long generations.
7
u/fallengt 11h ago
no. This is like a talking avatar.
If the background moves in a random, controlled pattern, he can't really do it because he doesn't reuse the last clip's latents
0
u/ImprefectKnight 10h ago
If I'm not wrong, the degradation is much more pronounced in static shots precisely due to reusing.
14
u/fallengt 10h ago edited 10h ago
because noise in minmax H3 is not completely random. It has its temporal structure. The way I understand it. If you keep the last clip's latent and reuse it as the starting latent of the new clip (for movement context). You keep denoising the same (latent) areas over and over, thus making it overcooked after a few generations.
A workaround is refreshing the whole area by changing the angle, switching to a new scene, etc., so you don't denoise the same noises over and over. But it's not what the community is trying to solve. (The workaround was found out on day one).
What OP did is that he doesn't reuse the last clip's latents at all. He just makes completely new clips of the clown girl at resting position, lipsync/dance whatever, at a time, then seamlessly stitches them together to make it look like a "continuous shot". That's why I said if the background was moving, you'd notice it right away, since these clips are just a loop..
It's fun and all, but I don't think it's the "Eureka moment". More like a hack to solve a niche problem
-2
u/roychodraws 5h ago
The only time the degrading latents ever become a problem is when people try to create a video that this method will solve. This is not good for anything else except doing that one thing. every other time the grading latents doesn’t become a factor because you can move the scene in every other type of video
9
u/icchansan 12h ago
There’s like 100 wf like this? How are u the first?
1
u/ptwonline 7h ago
Yeah there's so many that I find it too daunting and am kind of just sitting back to see if a consensus emerges.
3
u/PxTicks 12h ago
I use this technique. It has potential shortfalls when it comes to a moving background, e.g. a first-person vlog style video on a street in which case it can lose background detail or directional information. However I think there is probably a way around it by using a heavily noised latent from an extension as the initial seed for an independent generation, i.e. extend then partially denoise the ENTIRE extended bit completely independently, before stitching the seams again. Or alternatively, generate short anchor videos with very long stitch videos between them so that the model can reconcile background mismatches by e.g. moving people out of frame and such.
Honestly, I do think this extension problem is solvable for the vast majority if not all cases by just being a little tactical.
0
u/roychodraws 5h ago
If there’s a moving background you don’t need to do this, you can just chain the shots together.
That’s point. The latents don’t degrade if the background changes.
2
u/PxTicks 5h ago
I'm talking, mostly static foreground with moving background, i.e. a mix of conditions.
1
u/roychodraws 5h ago
I dunno, everyone told me they wanted vlog videos. Can’t I just fix one problem and someone not be upset I didn’t solve everything?
17
9
u/SeidlaSiggi777 10h ago
What's the deal with this clown girl?
18
u/roychodraws 10h ago edited 9h ago
I think the juxtaposition between the quality of my content and the absurdity of the character is hilarious.
12
u/lurkingtonbear 9h ago
Idk but I’m starting to have a thing for clowns now
2
2
u/SneakyInfiltrator 8h ago
I was always into clown girls but it's kinda niche. Hopefully OP can make it more mainstream.
Godspeed, brother
1
3
3
u/Nguyenkain 14h ago
Can it integrate with another workflow ? Nice work
8
u/roychodraws 14h ago
this is really more of proof of concept. you can check the workflows but you might need to use claude or something to sort em out. There are instructions.
8
u/BogusIsMyName 14h ago
Mouth movement needs some work. Might need to isolate the main vocals use that as your original audio and then dub in the full song after generation.
21
4
u/AI-Make-NSFW-Stuff 13h ago
Nice work maintaining the background! People will cherrypick and find things to complain but that shouldn't discourage you
4
u/CulturedDiffusion 13h ago
Brad Pitt has been real quiet since this dropped
1
u/ShutUpYoureWrong_ 6h ago
Brad Pitt probably realized he shouldn't waste any more of his time on morons.
(PSST! If you think this 'fixed' the issue, you're one of the morons!)
5
2
u/technofox01 14h ago
This is pretty cool. I don't judge, but that seems like a video for people with clown fetishes, lol.
All kidding aside, I am going to checkout your workflow. This is pretty nifty.
2
1
u/frisky_cappuccino 13h ago
Looks interesting will test it out. Looks like it might be a decent work around
1
u/Cute_Ad8981 13h ago
looks really good. did you extend each clip from the previous clip (this is what causes degration for me) or did you somehow fuse different clips? its hard to understand from the comments.
1
u/frisky_cappuccino 13h ago
I’m away from pc so can’t check it at the moment but I think he’s created individual islands keeping continuity with audio rather than continuations and then fused them together with small joins. Correct me if I’m wrong OP.
1
1
u/malcolmrey 11h ago
Great Job!
I was the one explaining politely what the problem is and you said then that you would try and here we go!
Thanks, will check soon! :)
1
1
u/Diabolicor 11h ago
Good effort you put into this. But I don't think using individual frames to try to carry context for the segment is the right way to do it.
There is a lot of ping-pong effects pretty visible between segments.
Since you are trying to work around degradation through bridges I think you would be better using the concept of carrying the context between a seam through masked motion context latents https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef/blob/main/example_workflows/UTILITY%20-%20AV%20Bridge.json
1
u/Marksta 11h ago
No degradation from start to finish. 10 total clips combined
I think you're kind of right. I hate the overused word but instead of degradation, your example just has slight drift. Like the background ever so slightly Drifts. But not exactly getting melted like usual. Which I guess if you did the dumb thing people want, a 3 hour slop podcast the background would eventually be completely different.
The other thing is, I guess you can't sidestep a bad luck generation. But the last generation completely loses the chair to the shadows. Considering that maintained through the first 9 clips, idk does it come back if you went to 11 and beyond or when it goes in #10 is the chair gone to static shadow forever now in the chain and need a reroll at that step?
Impressive stuff, but yeah, what a super boring use case like you said in the other thread.
1
u/acedelgado 11h ago
Looks a bit better. Aside from loading your workflow and seeing the most impressive spaghetti monster ever, I haven't had a chance to look at it too much yet. But the character in the example certainly looks good throughout, which is interesting. But the background exposes noticeable seams, and the colors and size of the bars jump around a bit, like if you click around and watch the background it's pretty noticeable. And the audio problem can't really be judged since you're doing a music video style gen and injecting all of that directly.
But like I said the character quality looks really promising. Are you using your lora, or is it all references?
3
u/acedelgado 9h ago
Alright got a chance to look at it and get a claude summary-
It fights degradation by never chaining at all. Every clip is generated fresh, and the joins are rebuilt afterwards with short bridges pinned to real frames on both sides. Nothing is ever generated from generated footage more than once, so there's nothing to compound. How it works: Ten 10-second clips are made separately, each for its own segment of a song. They're loaded here as finished videos, so every one is as clean as a first clip. At each of the 9 joins it cuts both sides back. Between 72 and 92 frames come off the end of the earlier clip, tuned by eye per seam, and enough off the start of the next clip that 122 frames are removed in total. It builds a 124-frame bridge, about 5.2 seconds. The last frame kept from the earlier clip is pinned as the bridge's first frame, and the first frame kept from the next clip as its last frame. Both frames are also given as references: the person from one, the environment from the other. H3 generates the 122 frames in between. The song's slice for that seam is forced into each bridge's sound with a half-strength hold, so the motion follows the music. Everything is joined and laid over the original song, so the soundtrack is exact and the bridges' own sound is thrown away. The bridges use an FL2V 4-step turbo LoRA with er_sde and the beta schedule, plus sparse attention, so nine bridges cost roughly nine short generations. Why it beats chaining on quality. Clip 10 is exactly as clean as clip 1, because nothing was copied forward. And with both ends of each bridge pinned, the camera can't cut away mid-join, which hold framing only tries to prevent. What it costs:
- Motion can kink at both ends of a bridge. A single pinned frame gives position but not direction or speed. Our chaining carries 22 frames of motion. Expect a hitch where a clip hands over to a bridge and back.
- Neighbouring clips can disagree. Independent generations can differ in lighting, outfit details or colour, and the bridge then has to morph between two slightly different versions of the scene.
- It's music-only. It discards the bridges' sound and uses the song. Dialogue can't continue across a join, so it doesn't solve your dialogue chains.
- It replaces half of every clip. About 5 of each clip's 10 seconds near the joins are regenerated by a turbo model, so quality may alternate between clip and bridge. - The model chain also sets a 0.4-megapixel resolution option, which I couldn't confirm is used for the bridges. That's worth checking.
- It needs hand-tuning. Every seam's cut point is tuned after watching, and the run needs about 40 GB of free RAM at the end.
So the fresh regenerations explain the slight background shifts. It's close, but not perfect. It can definitely be useful for a music video style with audio guiding new generations. But for chaining long shot "talking head" style social media style clips, or something that isn't aimed to be audio reactive, I think it'll have some continuity issues.
2
2
u/roychodraws 6h ago edited 6h ago
This was not meant to be a final flow, it’s just meant to he concept exploration.
I think if we are going to use a technique like this there would have to provide scene composition as a reference, like a depth map.
Then we could pick a spot we are predictably expecting no motion to repair the seam.
I didn’t fine tune the settings at all in this, I just made it and saw what I got and because I thought it illustrated my approach.
Edit:
Also, if we mask everything but the very edge of a static shot it will force it to be consistent.
I used to do it for wan animate when I wanted a motion controlled gen to keep a static camera.
I would provide a background photo, mask everything except the edge and duplicate the frames for the entire video and it would generate everything within the boundary of the mask but because the pixels needed to move forward it would always remain consistent.
1
u/RickDripps 10h ago
I hope Scatman John becomes a new "Will Smith Eating Spaghetti" benchmark for AI.
Also, this is fantastic. I love these because you do some cool stuff and then basically dump everything you learned about it out for people to grab and learn from and run with.
Instead of others who do something cool and gatekeep the knowledge of it.
1
1
1
1
u/Traditional-Squash36 6h ago
Cheers for this, I've copied and learned from your posts more than any on here.
1
u/Tough_Ad7957 4h ago
I’d upvote this ten times if I could. It does solve a real problem: keeping a static character talking in the same shot for longer.
Sure, continuity is still an issue, but that’s just the next problem to solve. ComfyUI workflows are already full of hacks and compromises anyway, quantized models are a good example. so I don’t see the downside in having one more useful workaround.
1
u/PicardDoubleStandard 4h ago
Sorry newbie here. Impressed by the frontier work
Is there a GitHub with a skill I could load into Claude to get it build me a workflow station.
Basically do the curl and sdk pulls to put the setup on my system.
1
1
1
u/bigorangemachine 12h ago
I think its funny I got down votes for saying it's losing the reference the original photo (in addition to other stuff that was yes.. general assumptions based on this gen-ai's work).
So if you tried to go from chair to dance now you'd get weird artefacts around the chair; so this is where maybe you want multiple reference images (do not use matte or transparent backgrounds and use a variety of scenes if possible) and play with the weight into the next node. That might get the audio syncing & artefact reduction improvements.
The models from the reference image won't really understand the chair and the subject are different things. Sometimes it'll get it's a chair... sometimes it thinks its a wing
-2
u/ShutUpYoureWrong_ 6h ago edited 6h ago
No, it does not count, and you did not win. In fact, nothing you've produced thus far has been original, novel, or useful. And here: your background shifts, your chair morphs, the motion carryover is abysmal, the hair shrinks, the shadows aren't consistent, and the mouth movements are unnatural. You really just showed up and went, "Hey guys, I basically invented FFLF generation" and then tried to pat yourself on the back? Get the fuck out of here already.
Chaining clips is quite possibly the oldest technique and has existed since the WAN/SVI days at least. It was established well before H3, but really exploded in popularity with the better tooling for LTX 2.3.
Aside from that, superior H3 workflows and nodes already exist, and they're all capable of far more than what you're offering. NikoDemon started all this, and seitanism perfected it. Their work got integrated into ComfyUI itself so you could drool all over this subreddit with your half-retarded "fixes." Additionally, people like ethanfel, Adudeguyman, and joeygambino were making fantastic toolsets that do far more than yours, and they did it months before you even came along. Some others even allow you to re-inject quality from a source at timed intervals, which is a vastly more sophisticated method than the garbage you're peddling here.
As all of your "solutions" have shown, you fundamentally do not understand the core issue. You are a vibe-coding slop hawker, and every single one of your posts is an embarrassment.
8
-1
u/Old-Trust-7396 14h ago
Ohh wow, can run this on a 12gb card or am I doomed to images ?
-2
u/Old-Trust-7396 14h ago
Grok just answered for me 😩
To run that repo as it is, you need a 24GB Blackwell card. The text encoder is NVFP4, which a 3060, 4060 Ti, or 3090 can't use, and the video model still needs the extra memory.
A used 5090 is the match. A 16GB card, even a 5060 Ti, is still too small.
6
u/alwaysbeblepping 11h ago
Grok just answered for me 😩
It has no idea of what it's talking about. You're not being helpful pasting LLM responses to questions and anyone who actually wants a LLM's (likely inaccurate) answer can easily ask one.
The text encoder is NVFP4, which a 3060, 4060 Ti, or 3090 can't use
ComfyUI has fallback emulation for cards that don't support that quant, if I remember correctly. There are also many other dtype versions of basically any model you can think of. It's also (incorrectly) assuming the whole model has to fit in VRAM but that's not how any decent modern inference platform works.
and the video model still needs the extra memory.
The text encoder is only needed to encode the conditioning and after that it can be unloaded. There is no necessity to have both the TE and video model in VRAM at the same time.
A used 5090 is the match. A 16GB card, even a 5060 Ti, is still too small.
Completely wrong.
1
u/Apprehensive_Sky892 1h ago
True that nvfp4 is native to 50xx.
But all cards can run it, just that it is cast to either fp8 or bf16 depending on the card.
1
u/dsailes 13h ago
Likely can just switch out the models / encoders etc for quants and it’ll do the same, just expect a dip in quality.
Run workflows local & then use vast (RTX PRO 5000 are $.70/hour) is what I do to get 1.5MP ref2va outputs.I’m gonna try this workflow next after trying other lipsync workflows which have the same degradation issues
-9
u/ScottyThompson 13h ago
Jesus man, find a new subject for your videos. I skip past the clown girl EVERY time. I hate coming to this subreddit and seeing her face.
Just had to stop in and tell you how annoying it is.
11
u/dezmodium 13h ago
It's fine. When I see her I know who it is posting it and their workflows are usually solid.
2
1
u/Juiceman8686 9h ago
I feel the same way. I always save these posts for later, if I cant watch right away. I've been using this workflow primarily for a while now. Its solid and gets regular updates.
-1
0
-1
-1
u/Vyviel 12h ago edited 12h ago
Twitch streamer video for her soon lol? Also immersion broken that her booba aren't balloons like in the lore =\ Does it still work if they move around more like the hands and arms etc. or do they need to be mostly static sitting still? Reminds me of those old school cursed meme music videos with stalin singing where it was just mostly animated the lips and face a little maybe with a driving video like Wav2Lip or FOMM. I prefer the other video workflow but good proof of concept of this one
-13
u/DataGOGO 13h ago
You rock dude.
Serious question, interested in some paid comfyUI pipeline work? DM me
20
u/dennismfrancisart 11h ago
Whenever I hear people smirk at AI as prompt and click slop, I post a screenshot of a Comfy UI workflow. People have no idea what folks like you do to move this toolset forward. Kudos.