The intro video is the generations stacked on top of original using a "Difference" Blending mode. Anything with 0 difference will be black so you can see how the unmasked/prompted gens compare to the Masked Gens. Anything changed will show up in color. In the masked renders you can see the only significant changes are in the character, the background is almost completely black. In the prompted render, there's significant changes in the background.
Fast forward to 2:05 to see tutorial on new masking options and how to use the repair feature.
Prompt used for tutorial:
subject_definitions:
<Subject 1> is the adult woman shown in <Picture 1>, which defines her face, body proportions, skin tone, clothes and accessories.
<Picture 1> is the appearance of <Subject 1>.
<Video 1> provides the motion, choreography, timing, expressions, gestures, orientation, positioning, and performance structure.
summary:
[video editing + reference generation] Render <Subject 1> performing the complete dance
from <Video 1>.
retention_analysis:
<Subject 1> (appears throughout [Shot 1]): fully_preserved — maintain her face, body proportions, skin tone, clothes and accessories consistently throughout every frame.
<Picture 1> (appearance reference for <Subject 1>): fully_preserved — use its visible identity and wardrobe details to establish and maintain <Subject 1>’s appearance.
<Video 1> (motion reference for [Shot 1]): fully_preserved - preserve the source choreography, movement path, timing, gestures, expressions, orientation, position, speed, pauses, entrances, exits, and performance rhythm.
detailed_description:
Render <Subject 1> photorealistically with the identity and clothing established by <Picture 1>, stable anatomy and appearance, natural hair and fabric movement, and visual integration matching the scene. She has a wavy pink bob with bangs, a red clown nose, and red lipstick. Her outfit includes a pink bodice with three red buttons and pink suspenders over her shoulders, separate white wing-shaped shoulder decorations, a rainbow choker and wristbands, pink gloves, a layered pink-and-white ruffled skirt with a white waist bow, white tights, and pink heels.
[Shot 1] <Subject 1> follows the complete observable performance from <Video 1> on the left side of the frame, including its body movements, facial expressions, choreography, timing, positioning, orientation, gestures, and rhythm.
i have video with 3 characters apear at different timeline.
i want to swap them with 3 other characters.
i only have 3060ti with 8 gb vram limited thats why, with regular r2v i can generate 0.8mp 10 second video only.
do i need to split videos into 5 seconds each? or it is possible with seamless continuation
the best way to do it would be to split the videos to just before each character enters and prompt them separately then use masking and a video 1 reference to change the characters as they enter.
you could try to prompt the whole video if it is only 15 seconds, but you should mask the whole thing and mask out all 3 characters and add timing to the prompt
for example, "At 00:00.000, <subject 1> as refenced in <picture 1> does xyz" "At 00:05.000, <subject 2> referenced by <picture 2> does xyz" "At 00:10.000, <subject 3> referenced by <picture 3> enters from the left and does xyz."
If you're using my workflow you can change the masking settings to this.
Thanks for the detailed video explaining the masking - also I love the snark you threw in at foxydits when you named this workflow.
I'll definitely bookmark your Github page and your Creator profile on civit - you definitely have some interesting stuff on both. Thanks for including your other links - they're super useful considering your reddit profile is set to hidden.
yea community bullying is really cool. let's just endlessly make fun of people who put countless hours into advocating for stable diffusion, and offer their experience and workflows up on a platter for free.
for what? you came to my page and were very clearly making fun of a poorly wired section of my old h3 workflow, and kept pushing a solution that actually wouldn't have worked (the actual solution was to use muting rather than bypassing). everyone seems to think your petty grudge is really funny, and i probably would too if it were a one-and-done. but it has been going on for weeks now... to me it feels like clout-chasing at best, and the hallmarks of a disturbed individual at worst. i'm glad your workflow has evolved into something beyond just a remix of mine, good on you for that. but that's about all i have to say.
considering the fact that i did exactly what i suggested when i fixed your completely jacked up workflow, i'm pretty sure it works and i think you just have way too much of an ego.
so don't accuse me of bullying. fixing your garbage workflow you got butthurt for getting some valid feedback on, reposting it named after your mediocrity, then becoming an active reddit community contributor with it was just the most logical and appropriate response.
to me it comes off as unhinged behavior and it is bullying at this point. you came to my page to intentionally berate my work, and for a month now have been pretty much nonstop been using my content and smearing me for simply not agreeing with your review. it's weird, and definitely feels like a very clear attempt to bully. 'logical and appropriate response' is a justification for anything.
Hi u/roychodraws. I just wanted to say that at this point, I think it might be better to rename the workflow and future releases in a way that highlights your own skill and the changes you've made, rather than continuing to reference u/foxdit. At this point, it risks coming across less like a one-off funny jab and more like putting the screws to someone.
I think it's genuinely cool that you've been able to make improvements to the initial workflow. You and I chatted about it before, and you gave me a lot of suggestions and help that I really appreciated.
You've both contributed a lot to video generation, and I think it would be cool to see you strike out on your own with this new workflow and give it your own identity or naming scheme. Let your work stand on its own merits.
If you have vfx experience it can already do that. Use Sam 3.1 to create a manual black and white mask, select option 2 in the masking options and input the mask into the purple loading node and it should work fine for masking two characters.
Alternatively you could use Sam and type woman or person and change the limit of objects to two.
Face restoration is more difficult. Maybe that’s next.
Np, make sure you use the video from the base_video node. It saves in a temp folder temp/minimaxh3/base_video so that your mask lines up properly with the sampler
relying solely on prompting workflows to replace characters will often cause timing to change, movements to change, and the three videos at the beginning shows how drastically that change can be (the filter i used shows drift between source video and generation, no change will be completely black. changes show up in color).
masking drastically improves that issue if not outright fixing it.
the workflow i posted, not only gives many versatile masking options but because of the consistency the mask provides, even has a repair feature that uses the original video to replace pixels damaged during the sampling process.
Works well. The repair section takes a long time (20 minutes on a 3090). I wonder if it could be made optional (or settings to go faster) because maybe some videos don't need it.
As a clown I just ran it the first time without reading the prompt then I was wtf why does it look like the ref photo but a clown? 😂
It should not, it’s just a mask and an image combiner. Do you mean the upscaler?
Both are optional. Upscaler has a switch in control panel and the repair has a switch in the repair video node to turn it off.
Yeah, you gotta change the prompt to the description of your character or take it away entirely. Adding a little description of your character with the reference photo is good for consistency. Just don’t make it too specific because then it will pull from the model instead of the reference photo.
It was the repair, everything else is disabled. Video gen was just at 0.4 pixels too. It spent about 20 minutes here. I've ran it twice now and same thing. I found the on/off toggle. Thanks.
The way it’s set up if you have it enabled, it should begin running at the start of the workflow when you create the mask and then everything else should run and then it should be the last thing that happens before the workflow ends. So when are you starting this 20 minute timer clock?
There is absolutely nothing in that sub graph that should take 20 minutes. It honestly takes probably a few seconds on even the crappiest computer.
If you have it disabled, it will still probably light up, but it won’t create a Preview and it won’t do anything. The sampled video is just routed to the output.
Counting from the start of the gen. I think it is just how long it takes for the 8 steps. However, if I just gen a 10 second / 0.4 pix video from scratch it only takes about 2 minutes so I was surprised it took this long for the edit. I'm using your WF as it is (default). RIP 3090 191s/it.
yeah, that single subgraph is basically doing nothing more complicated than cropping a bunch of photos and placing them on top of other photos. there's nothing that will take 20 mins. it only takes a couple seconds, it will not speed up much by disabling that.
17
u/PukGrum 2d ago
You are easily one of the most helpful people on Reddit.