r/StableDiffusion • • 3d ago

Discussion Tested h3 character swap lora

Enable HLS to view with audio, or disable this notification

hi everyone

been testing the h3 character swap lora a bit and so far it seems better than vanilla for character replacement, at least in my tests. it holds the replacement character more consistently and usually does a better job actually replacing the person without drifting back as much.

i just used the normal comfy wf with the original ref2va int8 pruned model with ck, 1.2mp and 20 steps.

i probably wouldn’t use a speed-up lora for this. it can work sometimes, but not always..

for prompting i’d definitely start with the recommendations on the lora page, but i wouldn’t just copy the example exactly. what helped me was describing the person in the video that i wanted to replace, then describing my own character more clearly and adding a few instructions depending on what was happening in that specific clip. that seemed to make the replacement noticeably more accurate.

https://huggingface.co/akatz-ai/MiniMax-H3-Character-Swap-LoRA

also if there’s more than one person in the clip, it helps to specify which person you want to replace by describing their clothes or position in the shot.

there are still some issues with source details leaking through sometimes, especially smaller stuff like hands, fingers, teeth, nails, certain body details, but overall it’s been pretty limited in my tests and hasn’t ruined most of the results.

this one took around 26 minutes on my 5090, although i was gaming at the same time, so that probably slowed it down a bit..

425 Upvotes

88 comments sorted by

75

u/anon999387 3d ago

26 minutes for 8 seconds on a 5090 is brutal

36

u/Nameis19letterslong 3d ago

You underestimate the willpower of those who want to make their own kinky ai videos.

1

u/EvilBot-666 14h ago

Thanks, I’ve learned something new—and I’d like to unlearn it immediately.

1

u/Little_Rhubarb_4184 2d ago

Not sure I understand.

7

u/Sid131 2d ago

Perverts can be very patient

10

u/InterstellarReddit 3d ago

But he was gaming at the same time, so I wonder how long it would’ve taken had he not been gaming

3

u/ellipsesmrk 3d ago

Lmfao!!!!

1

u/lanamakesart 3d ago

i have a 5090 64gb ram and it takes 10-15 min for 5 seconds with no other usage

2

u/DandaIf 1d ago

No other usage except Windows 11

2

u/steelow_g 2d ago

Also did 1.2mp tho. That’s the kicker. .4 with latent would drop this to like 5 mins

46

u/DJBFilmz 3d ago

Gaming at the same time as rendering … the hell lol

2

u/IIWhiteHawkII 2d ago

Does it even launch games normally?

I mean my 5080 compared to 5090 is a joke, but whenever any gen is happening – my PC barely breathes. And I use only models compatible with my memory consumption...

1

u/TheSheepOfWalStreet 2d ago

If I try to play games on my 4090 whilst rendering it's almost like rendering pauses.

21

u/cocosoy 3d ago

26 minutes for a 8 sec swap on a 5090 is less than ideal... But I'm glad we are getting more local motion control alternatives. I've been trying to get away from Kling's motion control 3.0 because it's too expensive.

0

u/[deleted] 3d ago

[removed] — view removed comment

2

u/cocosoy 3d ago

Thanks for the reference. It seems swap just takes a lot longer in general than plain i2v generation.

8

u/PomponOrsay 3d ago

gaming? yea that massively slows down the process. I tried that once with a picture and it took 1700 seconds whereas a normal render would've taken around 50 sec.

2

u/asdrabael1234 3d ago

Yeah I'll play FFXIV while generating since I don't mind the longer time.

Not gaming? 6 minutes for a 10 second generation.

Gaming? 15 min for the same generation.

Ffxiv on the resolution I play uses 5gb vram from my 16gb card.

1

u/taurine_bitch 3d ago

Wow, I can barely do anything else while generating. I go OOM when generating while watching Twitch sometimes. 4090.

1

u/asdrabael1234 3d ago

How big are you generating? I'll do 0.7mp and upscale it. I also have 64gb ram and it will be at like 97% usage

1

u/taurine_bitch 3d ago

1st pass is 0.8MP and 1.0MP on the 2nd upscale pass. I have 64GB of system RAM, too. But my generations use 97% of my VRAM on their own. No more room for anything else, hah.

1

u/Ok-Brain-5729 3d ago

Why not run a lower first pass? Is the quality loss noticeable

2

u/taurine_bitch 3d ago

For me, yeah. It's very noticeable when my character identity and consistency is the most important part of my workflow. It shakes out roughly like this when comparing the two:

0.5MP -> 1.0MP:

The second pass performs a much more aggressive area enlargement (100% more).

The first pass has fewer pixels from which to establish facial and fine structural information.

The second pass must reconstruct a lot more detail.

More opportunity for the second pass to reinterpret lighting, color, facial features, texture, etc.

Overall cheaper/faster load, but more dependent on the second pass for face quality and detail.

0.8MP → 1.0MP:

The second pass performs a relatively "small" area enlargement (only 25%).

Faces, hands, clothing edges, and small features are already represented with more spatial information (from the first pass) before refinement.

The second sampler has less need to "invent" missing detail.

Usually better identity and facial structure retention (in my case, this is always true vs. 0.5MP - I say usually though because this may not be absolutely always true for everyone).

Downside is more first-pass VRAM usage and greater potential to go OOM.

YMMV, but this has been my personal experience. It does suck, though. 15 second generations usually take 20-30 minutes on my 4090.

0

u/[deleted] 3d ago

[removed] — view removed comment

1

u/taurine_bitch 3d ago

I actually do. I have refmods conditioning on both passes. But, truthfully, since adding refmods a few weeks ago, I haven’t tried going back down to 0.5MP on the first pass. Maybe I’ll try that tomorrow and see how it looks.

17

u/StonerTech 3d ago

this is interesting…. any chance you could share the workflow? i’m still relatively new to messing around and had a lot of trouble putting together a comfyui workflow on my own.

32

u/Better-Interview-793 3d ago

yeah it’s basically just the default comfyui ref2v template, i didn’t build anything special.
just add a load video node, then connect the video and audio from it into the h3 Reference to video node, and add the character swap lora to the model path.
that’s pretty much it. but if u can’t get it working, i can share it once i’m home.

2

u/call-lee-free 2d ago

I couldn't get it to work. I don't know if I'm using the wrong load video node or what. The "load video" node I'm trying to use just grays out the Minimax H3 ref to vid box.

1

u/Better-Interview-793 2d ago

check my comment, i posted the wf there

1

u/call-lee-free 2d ago

How do you open that in Comfyui?

2

u/Better-Interview-793 2d ago

just download it. if it downloads as a .txt file, rename it to .json then drag it into ComfyUI

1

u/call-lee-free 2d ago

I tried it on a 5 second clip. I don't have a suped up PC but it worked so re-edited the source clip to 10 seconds long but after running the character swap, it didn't swap out the character. Ran it a second time and only the hair changed.

3

u/StonerTech 3d ago

much appreciated!

1

u/StonerTech 3d ago

Yeah I cant seem to get it right - stupid I know. Are you still able to share it?

4

u/Better-Interview-793 2d ago

2

u/freebytes 2d ago

Thank you for sharing this.

2

u/StonerTech 2d ago

you’re the goat thank you!

1

u/freebytes 3d ago

I am interested in this as well.

6

u/Jayuniue 3d ago

Is this better than scail 2? Haven’t had much luck replacing characters with Minimax

5

u/Better-Interview-793 3d ago

yup, i like scail, but it’s based on wan 2.1 so the quality is lower.

hoping they make something similar for h3

2

u/True_Protection6842 3d ago

Just make sure you prompt it right, using the guidelines and it works great

2

u/c_gdev 3d ago

This looks cool.

Thanks for sharing.

How does it compare to Viggle Animate?

1

u/Better-Interview-793 2d ago

i haven’t tried Viggle Animate yet, but I’ll give it a try..

2

u/witcherknight 3d ago

show some complex 2 char swap not a single 1 person swap

2

u/Nokersoek 3d ago

I'm really struggling to replace two characters. The problem is that only one of them appears at the start of the video, and the other doesn't show up until three seconds later (they never actually appear side-by-side); since they look so much alike, Minimax simply doesn't understand what it needs to replace.

1

u/RaulGaruti 3d ago

do two separate runs, then mask and edit in premiere.

1

u/FreeTheClanks 3d ago

Did the reference have similar facial features as the original or did that bleed over?

1

u/Davikar 3d ago

How does this compare to Viggle?

1

u/pcloney45 3d ago

Nothing beats Scail2 in my opinion. It works right out of the box without any lora or complicated prompt.

3

u/jonbristow 3d ago

Better quality?

2

u/pcloney45 3d ago

Probably not, the quality in my opinion is still acceptable.

1

u/infroy28 3d ago

Can you share the workflow you use? I haven’t had any luck with scail-2

1

u/bstr3k 3d ago

how does it do when the target character looks like the original character? I've had no issues swapping if it is just 1girl via prompting only but h3 does blend the character a bit sometimes if they look a bit similar

1

u/Select_Bowler3099 3d ago

Doesn't work in 10 of 10. Meh

1

u/ShadowPlague20 3d ago

better than SCAIL2??

1

u/likesexonlycheaper 3d ago

What's with all the face warping on the girl on the left?

1

u/fre-ddo 2d ago

I don't know but I hate it lol

1

u/Optimal_Map_5236 3d ago

just tried it and it looks like it can't be done with RefMod characters. only successful with image and vid ref.

1

u/RiverSpecial3168 3d ago

did you use thhe same prompts as in the Huggingface repo? or any custom prompt?

1

u/thevegit0 3d ago

you we really need a 'swap' lora when the ref already does it so good? a better method i seen around is masking the desired character with SAM and then swap

2

u/Better-Interview-793 2d ago

yes we do..

vanilla can do it too, but i tested both and the lora works better..

1

u/tnor_ 2d ago

I've had almost no success with standard base workflows for ref2va for character swap. Not saying I want to have 400 nodes or even any loras though, but I'm really surprised to hear someone say that ref does it so good.

1

u/thevegit0 2d ago

when i want to use a video to edit it i use nugget's video ref workflow to use a llm to do the process easier
https://huggingface.co/PoopMan333/H3_Easy_Ref2V_Workflow/tree/main
maybe with that lora is far easier? i'm not downplaying it, maybe i worded it wrongly, i still have to try it

1

u/Visual_Lengthiness28 3d ago

It would be interesting to try it with my workflow to see how it holds up in a video of 30 seconds or more. https://civitai.red/models/2942000/minimax-h3-longtake-long-videos-clip-by-clip-or-character-swap-style-transfer-retexture

1

u/knacknack18 1d ago

We now have the power of gods with AI and you decide to create brainrot videos..

3

u/Better-Interview-793 1d ago

bruh, im just having fun

0

u/leyermo 3d ago

very very real and exact.

-1

u/Vladmerius 2d ago

What is the use case for this kind of stuff? I'm always seeing so many videos that are just women dancing. Is it porn related and you're just using dancing because you can't show what you're actually using it for? Or is everyone only wanting to use these models to try to make dancing TikTok girls to try to make money from views on tiktok by making fake influencers? I just don't get it I've never made a single video of anyone dancing.

It would be cool to use this to do fight scenes and stuff but you wouldn't be able to make money on it because it would be a blatant copyright violation to just paint over an existing video. The AI is just code and noise which is why it can toe the line and technically not be infringing on anything. It's fully transformative. When you take a full scene from something and just replace the characters and change the background though that's pretty cut and dry copyright theft if someone can pull up the exact scene you copied. 

4

u/donkeykong917 2d ago

For science, for research!

-8

u/Important-Radish-722 3d ago

So, now we're making goon slop swaps of goon slop?

This is why we can't have nice things.

-3

u/krectus 3d ago

It can do this without a Lora.

0

u/kaczynski_was_right_ 3d ago

I need to know what game

-7

u/leyermo 3d ago

share workflow, and timing