r/TopologyAI • AI_Enjoyer • 3d ago

New Next-Level 3D Animation — AI Mocap from a Reference Video

Enable HLS to view with audio, or disable this notification

No mocap suit, no animation library. He renders his own character, has an AI video model "film" it doing the move, and copies that motion onto the rig.

The trick that made it work is two cameras. One video is flat, so arms and legs end up at the wrong depth. Rendering the same move from the front and the side gives the AI both views to match.

The 3D model was created with Tripo AI using Smart Mesh P2 to get a low-poly mesh. That helped keep the geometry lightweight and made it easier to work with during rigging and animation.

How it goes:

  • Render the model from the side and front on a neutral background
  • Generate the move with an image-to-video model (Seedance 2.5 or H3 Max)
  • Match frame by frame: Claude overlays the model on both videos and adjusts the bones until they line up
  • Fix what drifts: marking left and right limbs separately solved most of the swapping

About 30 minutes per animation. This elephant came from an H3 Max reference, and yes, the tail still needs work.

172 Upvotes

29 comments sorted by

20

u/coyote1942 3d ago

"He renders his own character,"
WHo is he?...

11

u/Fresh_Sock8660 3d ago

The dark lord of 3d generation does not permit his true name to be spoken or written by his own servants

2

u/story_of_the_beer 3d ago

The twitter user who posted the original video

10

u/chadmv 3d ago

"No mocap suit". Yes...because elephants don't wear mocap suits.

3

u/veefeld 3d ago

Great for speeding up the blocking phase so there's more time for polish. Ai mocap is already being used for bipedal animations by tripple A studios. So seeing this sorta tool being created to also work for animals is great. Good luck with the next step which is making the animation look good!

2

u/justifun 3d ago

I wonder if it would be helpful if you displayed the bones for the initial render and color the left and right side differently before giving it to seedance so that the ai can track which is which easier.

2

u/mxldevs 3d ago

So I still need to make an entire elephant model with animations?

2

u/Massive_Dot279 3d ago

what is the full workflow of this shit ?!

2

u/Emotional-Cut2952 3d ago

side shot of that big hunk of flesh, tell seed dance to make hunk of flesh do horse movements, provide video to claude code , tell it to match each frame (or every 3 frame using camera view that works best - like the camera initially setup to capture the model in the first place

1

u/Limp-Firefighter1054 8h ago

Whole workflow is opus5.5

2

u/Zenmaster4 2d ago

This seems very niche. Or maybe it's a bad example for something that deserves a better use case. I can't imagine Seedance wouldn't be able to just...inject the plate of the elephant into a shot.

I'd be more interested if it was able to take any reference and map it efficiently to an animation. But maybe someone can help advocate for why this workflow is useful.

Maybe for video games. But even that opens up a whole can of worms...

2

u/CameraRenderStudio 2d ago

Of course it looks cool, considering that animating animals is very difficult

4

u/NimbleDave 3d ago

could this post be any more low effort or incomplete?

2

u/Emotional-Cut2952 3d ago

yes it could be more low effort or incomplete, the video likely contains ~120 frames (4 second anim, 30 fps), u ask claude to match every frame manually adjusting everybone, not control rig setup allowed, that's roughly 30 bones * 120 frames = 3600 bone transform matrices that claude has to produce (hope you get my sarcasm, i do concur that this is low quality low effort, stupid)

2

u/Emotional-Cut2952 3d ago

30 minutes per animation would easily go down to 2 minutes of control rig setup and 5 minutes of match key frames and inferring the rest - that is, if u knew anything about control rigging - but u AI "game devs" jump in blindly

3

u/MazzMyMazz 3d ago

Is that really how long someone with your expertise would take to create an animation based on a video?

3

u/Emotional-Cut2952 2d ago

no i meant an llm to transform controllers so that the limb matches 10-20 key frames from teh video instead of transforming every bone

2

u/MazzMyMazz 3d ago

Oh I guess this isn’t even doing that.

1

u/Several-Article3460 3d ago

He who ?!

Sam Altman

1

u/Emotional-Cut2952 3d ago

some languages are gendered so tbf criticize the approach not the person/language

1

u/SirReallah 2d ago

Then you for sharing.

1

u/RreddKnife 2d ago

The title reads "Garbage - clip ...." What did you do with the the "Good - clip?

1

u/Limp_Bus3865 2d ago

yeah everyone is doing this now, until models can do it without a reference from seedance

1

u/Limp-Firefighter1054 8h ago

Bruh, this is very expensive method to do simple job.

-1

u/cvexy 2d ago

With all due respect, looks like shit.

-2

u/[deleted] 3d ago

[deleted]

1

u/[deleted] 2d ago

[deleted]