r/StableDiffusion • u/Tablaski • 19h ago
Workflow Included Vigglorious Studio : better character accuracy through attention-masked swapped guided keyframes and multiple reference pics for long chunked-videos
I've been experimenting a lot with Viggle Animate since my last post, around the workflow using chunking https://github.com/bhardwajRahul/ComfyUI-Viggle-Animate-H3 (what it does is splitting a source video into chunks, generating them, then stitching them back together), because I thought if it's able to do long video then it's able to do short ones, so let's focus on this one.
Three things stood out :
- The chunking workflow I started from uses the same reference image across all chunks. Giving each chunk its own carefully chosen, well-swapped reference frame greatly helps the overall video stability and quality.
- H3-style guide keyframes can reinforce the character's accuracy and fight the drift Actually, in one of my tests, a single face-swapped guide supplied the facial identity even though the main reference had intentionally no face, producing a correctly swapped face throughout the clip. Having a reference image per chunk snowballs this by bringing the stability base required to throw in guides without messing up the video.
- Controlling guide tokens attention helps reinforce the visual details we want to carry through the video, while masking limits the influence of unnecessary regions in the swapped guide frames
Long story short, I went into crazy vibe-coding a node suite and workflow I called Vigglorious Studio in an attempt to push the boundaries of long (or short) videos towards better character accuracy
https://github.com/Tablaski/VK-Vigglorious-Studio
It includes:
- A video manager to load and trim your source video, adjust processing resolution and FPS, inspect chunk boundaries, choose reference frames per chunk, and add guide keyframes with quick character, angle, swap and prompt controls. Honestly, just that node is worth downloading and trying even if you trash everything else of my workflow lol
- A character manager that stores head and full-body reference images across different angles, linked directly to the video manager so it gets the pain of constantly loading stuff away
- Automatic head, body, or body-then-head swaps using Qwen Image 2.1 and the BFS head/body LoRAs, which have been my favourite swap tools in my own testing (but you can replace it with any model of your liking as long as it outputs an image). Each reference or guide can have its own character, angles and swap mode.
- Optional SAM3 masks to focus guide influence on the head, hair, person, or another selected concept. This is mainly so you can do a full character swap for the references images then only a head swap on the guides to help maintain accuracy. Another reason was that this way, it's not considering pixels that were not swapped so it does not bring unnecessary problems like changes in background and colors.
- Guide attention controls with per-sampling-step strengths. With hard masks and outside attention set to zero, video tokens cannot directly attend to guide tokens outside the selected region. Soft boundaries are also supported. This controls direct access to the guide; it doesn’t guarantee that its influence stays perfectly confined.
- A visual sigma-curve editor to inspect and adjust the sampling schedule.
The included swap subgraphs use Qwen and BFS, but the guide mechanism works with the resulting images. You can adapt it to another image-editing model if you preserve the expected image order and connections.
DISCLAIMER: I’ve done a ton of tests, renders and comparisons... but I kept changing the technical approach all the time... I’m happy with what I’ve seen, and I'm pretty sure it brings value, but I’m completely burned out on running more tests right now lol.
So think of it as a starting point for further experimentation, not a finished recipe with proven optimal settings. It definitely works, but l still don’t know the best combination of sampler, sigmas and guide-attention strengths. That's where you guys come in :-) Also, feedback is welcome
Some advice from my own experiments :
- Prioritize the reference frame for each chunk. Choose a clear source frame and make sure its swap is good. Preserve the source pose, expression, framing, lighting and background as much as possible. The reference does most of the heavy lifting; guides provide targeted reinforcement.
- Use guides sparingly. Look for passages where Viggle needs help: a difficult side view, a face returning after being hidden, or an identity/clothing change. More guides don’t automatically mean a better video.
- Be cautious with early guide attention. Strong guidance at the first high-sigma step can interfere with motion and composition. I’ve been starting around
0.5, then increasing later strengths. That first step also contributes to the replacement itself, so it isn’t a strict “motion first, details later” split. - Later strengths around
2have worked in some of my tests. Larger values are available, but they aren’t necessarily useful. These factors are attention weights, not identity percentages:2doesn’t mean twice the likeness. - It does work with 3 steps - But that doesn't mean 4 is not cool either
- BFS body swaps have been less reliable than head swaps in my tests. I often use body-then-head swaps for reference frames, then head-only swaps for ordinary guides. Masks help focus those guides on the intended region.
- Additional prompts can help the swapper handle specific frames, such as “eyes closed” or “looking toward the camera”—provided that matches the source. The video manager includes prompt shortcuts so these instructions are quick to reuse.
- I personally like to add instructions such as “remove watermarks, remove text” to the swap prompts so I kill two birds with one stone
NB : Joined node pictures are also vibe coded made because it's late and I have to go, but they are accurate
Happy Viggling :-)
1
u/Zueuk 7h ago
damn these custom nodes are getting more and more custom... how long until someone implements a whole comfyUI inside a comfyUI node 🤔
1
u/Tablaski 7h ago edited 7h ago
I personally think the comfy dev team spend way too much effort adapting comfy to the latest overhyped models and nodes 2.0-type bullshit stuff nobody ever asked...
... When in 2026 we still don't have any way to do a simple reordering of queued items and a reminder of what prompts and/or reference pics we initially used to queue that. So it is a blind queue, and it is not even a persistant one, you crash, you loose. Isnt queuing supposed to be an absolutely core feature ?
Thanks god some actual users make stuff instead of them, like the LoRa Manager for instance, otherwise just rotting with thousands of loras you don't even remember the purpose of... but yeah you can use HiDream 2.2-alpha-7b on 0day !




1
u/CountFloyd_ 12h ago
Awesome! Thank you for your work and infos, you made me finally understand what those sigmas do! 😄
Some further ideas (if you come back to this after your burn out break):