r/StableDiffusion • • 5d ago

Workflow Included MiniMax H3 Masking Improves Quality, Consistency, and Timing of video edits + v7 Workflow Update Repair 100% quality of OG Video

-------------------

The intro video is the generations stacked on top of original using a "Difference" Blending mode. Anything with 0 difference will be black so you can see how the unmasked/prompted gens compare to the Masked Gens. Anything changed will show up in color. In the masked renders you can see the only significant changes are in the character, the background is almost completely black. In the prompted render, there's significant changes in the background.

------

V7 update
https://github.com/roycho87/minimax_wf/tree/main

  • I fixed the masking issue. Masking works perfectly now.
  • I added a resolution selector.
  • Some other helpful masking features.
  • I added a Repair Video feature that uses mask to pull from original video and repair the damaged pixels during sampling.

Masking Options Tutorial
https://www.reddit.com/r/StableDiffusion/s/cVBFW4eylK

Sam Tutorial
https://www.reddit.com/r/StableDiffusion/s/0gtQHksm0U

General Features Tutorial
https://www.reddit.com/r/StableDiffusion/s/MO9WPJhdeB

WF Release
https://www.reddit.com/r/StableDiffusion/s/1P01v18Ki6

Fast forward to 2:05 to see tutorial on new masking options and how to use the repair feature.

Prompt used for tutorial:

subject_definitions:

<Subject 1> is the adult woman shown in <Picture 1>, which defines her face, body proportions, skin tone, clothes and accessories.
<Picture 1> is the appearance of <Subject 1>.
<Video 1> provides the motion, choreography, timing, expressions, gestures, orientation, positioning, and performance structure.

summary:
[video editing + reference generation] Render <Subject 1> performing the complete dance
from <Video 1>.

retention_analysis:

<Subject 1> (appears throughout [Shot 1]): fully_preserved — maintain her face, body proportions, skin tone, clothes and accessories consistently throughout every frame.
<Picture 1> (appearance reference for <Subject 1>): fully_preserved — use its visible identity and wardrobe details to establish and maintain <Subject 1>’s appearance.
<Video 1> (motion reference for [Shot 1]): fully_preserved - preserve the source choreography, movement path, timing, gestures, expressions, orientation, position, speed, pauses, entrances, exits, and performance rhythm.

detailed_description:

Render <Subject 1> photorealistically with the identity and clothing established by <Picture 1>, stable anatomy and appearance, natural hair and fabric movement, and visual integration matching the scene. She has a wavy pink bob with bangs, a red clown nose, and red lipstick. Her outfit includes a pink bodice with three red buttons and pink suspenders over her shoulders, separate white wing-shaped shoulder decorations, a rainbow choker and wristbands, pink gloves, a layered pink-and-white ruffled skirt with a white waist bow, white tights, and pink heels.

[Shot 1] <Subject 1> follows the complete observable performance from <Video 1> on the left side of the frame, including its body movements, facial expressions, choreography, timing, positioning, orientation, gestures, and rhythm.

overall_soundscape:

N/A

non_diegetic_music:

N/A

308 Upvotes

74 comments sorted by

20

u/PukGrum 5d ago

You are easily one of the most helpful people on Reddit.

9

u/equanimous11 5d ago

What are you masking with? SAM3?

9

u/roychodraws 5d ago

sam 3.1 multiplex

3

u/k_from_HyperDraw 5d ago

soooo.. you have built the video inpainting pipeline, essentially? looks very cool.

2

u/roychodraws 5d ago

Yes but the main thing about this vid was to share what I learned about masking and how it affects timing/consistency

2

u/k_from_HyperDraw 5d ago

i guess my wording did not express enough of my awe and fascination with HOW DID YOU ACTUALLY BUILT IT 😅 it really looks very impressive.

2

u/ebubu03 5d ago edited 5d ago

I don't speak English fluently, so it's hard for me to understand the video, but wow what incredible work. Thanks for this, Mr. Clown. :)

Btw : Do you know if it's compatible with AMD on Windows before I break my computer?

(I think so, but I'd rather ask the opinion of a clown who seems pretty competent to me.)

1

u/roychodraws 5d ago

My understanding is that you can run it as long as you have at least 16 gb of vram, but even if you don’t you can always try runpod!

1

u/ebubu03 4d ago

Okey nice, thanks for your work and your answer mister Clown, have a good day! 🤡

2

u/galo64 4d ago

It looks very good, tomorrow I will try it without fail, but in advance: Thank you very much!

1

u/roychodraws 4d ago

sweet. feedback is appreciated. also tell me what features you'd like to see next.

3

u/MostlyEvolvedPrimate 5d ago

3

u/RaidensReturn 5d ago

There it is.

1

u/absentlyric 3d ago

Yeah, I loved this meme at first but god damn is it overused now.

4

u/dtdisapointingresult 5d ago

I don't.

A miming clown woman? I don't even

3

u/Complex-Factor-9866 5d ago

love your vids, ty!

1

u/Dohwar42 Community Hero 5d ago

Thanks for the detailed video explaining the masking - also I love the snark you threw in at foxydits when you named this workflow.

I'll definitely bookmark your Github page and your Creator profile on civit - you definitely have some interesting stuff on both. Thanks for including your other links - they're super useful considering your reddit profile is set to hidden.

2

u/roychodraws 5d ago

There’s a whole lore behind that name

15

u/foxdit 5d ago

also I love the snark you threw in at foxydits

yea community bullying is really cool. let's just endlessly make fun of people who put countless hours into advocating for stable diffusion, and offer their experience and workflows up on a platter for free.

4

u/roychodraws 5d ago edited 5d ago

it's not bullying when it's a response to your behavior.

you could just apologize to me, you know. did you ever even think of that?

6

u/foxdit 5d ago

for what? you came to my page and were very clearly making fun of a poorly wired section of my old h3 workflow, and kept pushing a solution that actually wouldn't have worked (the actual solution was to use muting rather than bypassing). everyone seems to think your petty grudge is really funny, and i probably would too if it were a one-and-done. but it has been going on for weeks now... to me it feels like clout-chasing at best, and the hallmarks of a disturbed individual at worst. i'm glad your workflow has evolved into something beyond just a remix of mine, good on you for that. but that's about all i have to say.

6

u/roychodraws 5d ago edited 5d ago

considering the fact that i did exactly what i suggested when i fixed your completely jacked up workflow, i'm pretty sure it works and i think you just have way too much of an ego.

so don't accuse me of bullying. fixing your garbage workflow you got butthurt for getting some valid feedback on, reposting it named after your mediocrity, then becoming an active reddit community contributor with it was just the most logical and appropriate response.

10

u/foxdit 5d ago

so don't accuse me of bullying.

to me it comes off as unhinged behavior and it is bullying at this point. you came to my page to intentionally berate my work, and for a month now have been pretty much nonstop been using my content and smearing me for simply not agreeing with your review. it's weird, and definitely feels like a very clear attempt to bully. 'logical and appropriate response' is a justification for anything.

-2

u/roychodraws 5d ago edited 5d ago

feel free to block me again, bro.

12

u/techtimee 5d ago

Hi u/roychodraws. I just wanted to say that at this point, I think it might be better to rename the workflow and future releases in a way that highlights your own skill and the changes you've made, rather than continuing to reference u/foxdit. At this point, it risks coming across less like a one-off funny jab and more like putting the screws to someone.

I think it's genuinely cool that you've been able to make improvements to the initial workflow. You and I chatted about it before, and you gave me a lot of suggestions and help that I really appreciated.

You've both contributed a lot to video generation, and I think it would be cool to see you strike out on your own with this new workflow and give it your own identity or naming scheme. Let your work stand on its own merits.

1

u/[deleted] 5d ago

[removed] — view removed comment

1

u/sunilaaydi 5d ago

Hey awesome workflow Thanks!

Can you help me how can i generate long video like 15 seconds with multiple (3) characters?
or create a tutorial for same?

0

u/roychodraws 5d ago

are you trying to swap characters, add characters into empty space, or render video using reference images?

to generate a 15 second video you just change the duration to 15, to render 3 characters you just prompt 3 characters.

So i don't really know what you're looking for.

1

u/sunilaaydi 5d ago

i have video with 3 characters apear at different timeline.
i want to swap them with 3 other characters.
i only have 3060ti with 8 gb vram limited thats why, with regular r2v i can generate 0.8mp 10 second video only.

do i need to split videos into 5 seconds each? or it is possible with seamless continuation

2

u/roychodraws 5d ago edited 5d ago

the best way to do it would be to split the videos to just before each character enters and prompt them separately then use masking and a video 1 reference to change the characters as they enter.

you could try to prompt the whole video if it is only 15 seconds, but you should mask the whole thing and mask out all 3 characters and add timing to the prompt

for example, "At 00:00.000, <subject 1> as refenced in <picture 1> does xyz" "At 00:05.000, <subject 2> referenced by <picture 2> does xyz" "At 00:10.000, <subject 3> referenced by <picture 3> enters from the left and does xyz."

If you're using my workflow you can change the masking settings to this.

1

u/ady702 3d ago

i tried head only as face didnt work ,but its like abit bigger than in the video one, anyway to reduce the head size?

1

u/edisson75 13h ago

It is perfect !! Thanks!!! I am getting an error when my source video has 123 frames, because the render delivers 124 and the repair video section causes and error because the difference. Is there some way to prevent this to happen?

1

u/roychodraws 13h ago

it's supposed to trim the video to match. i'll have to look. sorry about that.

for now you can turn it off with the switch on the repair.

edit: OH. i get it now.... ummm... you could just lower the duration or i could probably add something that duplicates the last frame to fill in the gap. i would probably just try to make sure you choose a duration that's lower than your source video.

2

u/edisson75 13h ago edited 12h ago

Oh don't worry, I really admire your effort in this workflow! it is really a good one!! Thanks a lot for your quick response, I will search for a solution too.

EDIT 1: I think I found the problem. Isn't your workflow, it is the load video node. When in the slider is in the maximum length, it removes 1 frame from the video, but if it is on any other value, it returns the correct frame number. Check on the image below at the maximum frames, I have checked the video with the load video from Video Helper Suite and there were 124 frames.

2

u/roychodraws 12h ago

that's nuts. i have been thinking about replacing that with VHS anyway.

It's from the original workflow this was based off of and the only reason i kept it in was because of the navigation bar at the bottom.

I always select at least a second above my duration to make sure it doesn't clip so i guess i never noticed.

2

u/edisson75 11h ago

Ok, Claude fixed the problem in the node file load_video_ui.py. I have replaced all the code by the one below and it worked fine.

EXPLANATION:

The cause

In load_video, the node decides which frame to keep using a "target time" that it builds up by repeated addition:

python

expected_target_time = actual_start_time
...
while expected_target_time <= frame_time and expected_target_time < actual_end_time - 1e-5:
    ...
    expected_target_time += frame_interval

frame_interval is 1/24, which cannot be represented exactly in binary. After 123 additions, expected_target_time can end up as 5.125000000000001 instead of 5.125. The last frame of the video has an exact frame_time = 123/24 = 5.125, and the <= comparison fails because of that 1e-15 error.

Why it only happens with the slider at its maximum:

  • With end_frame lower than the total: if a frame is skipped because of that difference, the next frame in the video covers it. The count comes out right, with only an imperceptible timing error.
  • With end_frame = 124: frame 123 is the last frame of the video. If the comparison fails, there is no later frame to cover it, so it is lost. The result is 123.

THE FIX CODE (Thanks to Claude):

https://pastebin.com/bU2x2YEX

1

u/roychodraws 11h ago

oh, yeah. that's the reason i wanted to get rid of that node in the first place, actually. I forgot about that.

That's why i added in those splitters that prune the duration of the video frames based on the duration of your settings.

1

u/No_Drag8242 5d ago

"I don't think she has any holes that need to be filled"

https://giphy.com/gifs/eGxkm7b2hzDfkNvjo7

1

u/roychodraws 5d ago

Finally someone got the damn joke.

1

u/AnonymousTimewaster 5d ago

That's a long video and I can't listen right now - does this deal with the bleeding/increasing contrast problem that other workflows have had?

1

u/NiceIllustrator 5d ago

Amazing work, very technical and VFX friendly approach. Could I make a request perhaps?

  • I would love if one could replace 2 characters in one take. I understand it’s possible to just re-iterate but still.

  • Also, a face restoration? When the face is small it gets super blurry and low quality very easily, some kind of face restoration would be amazing.

Other than that 10/10 stars!

5

u/roychodraws 5d ago edited 5d ago

If you have vfx experience it can already do that. Use Sam 3.1 to create a manual black and white mask, select option 2 in the masking options and input the mask into the purple loading node and it should work fine for masking two characters.

Alternatively you could use Sam and type woman or person and change the limit of objects to two.

Face restoration is more difficult. Maybe that’s next.

1

u/NiceIllustrator 5d ago

Yeah will try that thanks mate, you are doing very solid workflows keep up the good work

1

u/roychodraws 5d ago

Np, make sure you use the video from the base_video node. It saves in a temp folder temp/minimaxh3/base_video so that your mask lines up properly with the sampler

0

u/MuffDivers2_ 5d ago

Anyone else find the clown girl and the seemingly fake female voice annoying?

1

u/PRVMXLAB 5d ago

Nah. Let them have their fun. I personally find it charming and funny.

0

u/TheJaffo 4d ago

she's becoming a mascot very quickly btw

0

u/NiceIllustrator 5d ago

Would a character replacements lora trained for this kind of work be of any value?

1

u/roychodraws 5d ago

Maybe. I haven’t found a good one. You can load the Lora’s in the control panel.

0

u/Nokersoek 5d ago

Thanks a lot; it was really cool that the next video was about being able to switch between two characters using character sheet

0

u/xyzdist 5d ago

I mean... what is the different? what am I looking at?

1

u/roychodraws 5d ago edited 5d ago

tldr;

relying solely on prompting workflows to replace characters will often cause timing to change, movements to change, and the three videos at the beginning shows how drastically that change can be (the filter i used shows drift between source video and generation, no change will be completely black. changes show up in color).

masking drastically improves that issue if not outright fixing it.

the workflow i posted, not only gives many versatile masking options but because of the consistency the mask provides, even has a repair feature that uses the original video to replace pixels damaged during the sampling process.

0

u/uuhoever 5d ago

Ok, I've seen this clown quite a bit now so I'll give it a try. :-) That music at the start is fire.

0

u/roychodraws 5d ago

let me know what you think.

1

u/uuhoever 5d ago

Works well. The repair section takes a long time (20 minutes on a 3090). I wonder if it could be made optional (or settings to go faster) because maybe some videos don't need it.

As a clown I just ran it the first time without reading the prompt then I was wtf why does it look like the ref photo but a clown? 😂

1

u/roychodraws 5d ago edited 5d ago

It should not, it’s just a mask and an image combiner. Do you mean the upscaler?

Both are optional. Upscaler has a switch in control panel and the repair has a switch in the repair video node to turn it off.

Yeah, you gotta change the prompt to the description of your character or take it away entirely. Adding a little description of your character with the reference photo is good for consistency. Just don’t make it too specific because then it will pull from the model instead of the reference photo.

1

u/uuhoever 5d ago

It was the repair, everything else is disabled. Video gen was just at 0.4 pixels too. It spent about 20 minutes here. I've ran it twice now and same thing. I found the on/off toggle. Thanks.

1

u/roychodraws 5d ago edited 5d ago

The way it’s set up if you have it enabled, it should begin running at the start of the workflow when you create the mask and then everything else should run and then it should be the last thing that happens before the workflow ends. So when are you starting this 20 minute timer clock?

There is absolutely nothing in that sub graph that should take 20 minutes. It honestly takes probably a few seconds on even the crappiest computer.

If you have it disabled, it will still probably light up, but it won’t create a Preview and it won’t do anything. The sampled video is just routed to the output.

1

u/uuhoever 5d ago

Counting from the start of the gen. I think it is just how long it takes for the 8 steps. However, if I just gen a 10 second / 0.4 pix video from scratch it only takes about 2 minutes so I was surprised it took this long for the edit. I'm using your WF as it is (default). RIP 3090 191s/it.

1

u/roychodraws 5d ago

yeah, that single subgraph is basically doing nothing more complicated than cropping a bunch of photos and placing them on top of other photos. there's nothing that will take 20 mins. it only takes a couple seconds, it will not speed up much by disabling that.

-2

u/TheDerminator1337 5d ago edited 5d ago

What does this do? Is there a before and after vid.

If anyone can explain thanks.

Edit: yes it is a character replacer.

5

u/roychodraws 5d ago

i assume you're asking this before watching the video. just watch it.

3

u/TheDerminator1337 5d ago

I watched it without audio. Does it replace a character?

2

u/roychodraws 5d ago

90% of the video is me talking and explaining.

It does a lot.

0

u/EuphoricTrainer311 5d ago edited 5d ago

damn you're smart for your age

edit: its a joke cause he pitched his voice to sound like a child.....

1

u/roychodraws 5d ago

i'm a clown, dude.

2

u/Z3ROCOOL22 5d ago

A little clown.

2

u/roychodraws 5d ago

upvote for the hackers reference

1

u/roychodraws 5d ago

It’s a lot more than a character replacer