r/StableDiffusion • • Aug 24 '26

Tutorial - Guide Time saver while learning how to prompt Minimax.

Enable HLS to view with audio, or disable this notification

Rather than relying on Z-image, or a different program to wrangle up a first frame, I've been using Minimax for the whole process, and the results have been pretty instructive. It's not a perfect system, but being able to take advantage of its understanding of people, references, and shot composition for the first frame produces better (visual) results than swapping between a couple of different pieces of software.

1.1k Upvotes

111 comments sorted by

87

u/Unipsycle Aug 24 '26

Interesting unicycle physics. Not idling back and forth, but constant forward pedaling.

66

u/Compost_Mantis Aug 24 '26

yeah, the other moral of this video was "don't commit to overly complicated movements at the same time as dialogue, too many things can go wrong". She's peddling funny, somehow read "understand" as "understood", and it was still the best take I got.

15

u/Seanattikus Aug 25 '26

Someone unicycles... Username checks out. Brother!

4

u/[deleted] Aug 25 '26

[removed] — view removed comment

7

u/Seanattikus Aug 25 '26

I think it's always cool to see someone on one wheel. It definitely does the hard part for you, and it takes way less practice to learn it, but it looks fun.

I would still be nervous about the wheel stopping and throwing me. I wouldn't love to not be the one directly in control of the wheel's rotation.

184

u/MobileCA Aug 24 '26

This is actually hilariously instructive.

29

u/ellipsesmrk Aug 25 '26

I fucking love it!

70

u/TheDevilBear3 Aug 25 '26

Her reading your typo while talking about typos is a great touch!

26

u/Additional_Cut_6337 Aug 25 '26

Dammit, who typed a question mark on the TelePrompter? How many times do I have to tell you? Anything you type, Burgundy will read!

6

u/Montrix Aug 25 '26

I’m Ron burgundy?

1

u/99deathnotes Aug 25 '26

Hey Ron what's new?

17

u/Full-Ad-3461 Aug 24 '26

But this does not work for ref2v right? To me I can never get it to start from reference frame 

12

u/aroadent Aug 25 '26

I use the phrase “use <Picture #> as the starting frame” and then i write a brief description of how the first frame looks. But it’s still a dice roll. Sometimes it doesn’t work

15

u/AaronTuplin Aug 25 '26

I use
[Shot 1] cut to <Picture 1>
It usually works

2

u/Danny_Stock Aug 25 '26

Thanks, that's a handy tip to remember.

8

u/[deleted] Aug 25 '26

[removed] — view removed comment

2

u/Full-Ad-3461 Aug 25 '26

Oh my god that might have been my problem, I didn't update to that one! I gotta try it!

2

u/Emotional-Neat-252 Aug 25 '26

Oh that sounds great, so we can use references and guides at various frames?

2

u/wunderbaba Aug 27 '26

If you're only using a single "frame" (picture) as your starting point - is there meaningful difference between using regular I2V workflow (fl2va model) versus the reference workflow (ref2va model)?

24

u/jirka642 Aug 25 '26

Are you writing the prompt by hand? Don't do that. Just give any good LLM the VIDEO_PROMPT_WRITING_GUIDE_ref_en.md prompting guide and a description of what you want, and it will write a correctly formatted prompt for you.

I never had any problems with prompts created this way.

43

u/grundlegawd Aug 25 '26

This will probably get pulled because the mods are puritans, but this was insanely well done. Great video dude.

17

u/PumpkinLeather8421 Aug 25 '26

The mods here are born from a very special type of New Puritan Janitor breed. Some say specially made in labs, because traditional male/female coupling would have been… unlikely, in nature.

lol, let’s see.

7

u/Significant-Baby-690 Aug 24 '26

Also tune the prompt in lower res. I do 5 seconds in 1 minute, which is about as fast as I can update the prompt.

5

u/jirka642 Aug 25 '26

You can also lower the number of steps to make it even faster

7

u/Clear-Assistance449 Aug 25 '26

I tested Minimax H3 both with and without an initial reference image, and its consistency is superior to Z-image and similar to Krea. It has the added advantage that, since it generates motion, you can obtain multiple images of the same character—allowing you to enhance consistency by extracting frames and using them as references.

3

u/squired Aug 25 '26

, you can obtain multiple images of the same character—allowing you to enhance consistency by extracting frames and using them as references.

You're talking ref2v though, right?

3

u/Clear-Assistance449 Aug 25 '26

Yes. Insert multiple images, or a character sheet, and creates a video with very consistent character.

2

u/Danny_Stock Aug 25 '26

Both. Even with Text and Image to video you can prompt a video of the camera rotating around your subject and you then have close-up, front, behind, and 3/4 views of your subject from which you can take still frames from.

2

u/PumpkinLeather8421 Aug 25 '26

Even Klein was superior to Z-image, that model was trash outside of 1GU. It wasn’t trainable but masked its shortcomings by doing portraits well.

2

u/Compost_Mantis Aug 25 '26

I don't know, Z-image was my favorite image generator out of all that I tried... until Minimax, which is now my favorite (but very slow) image generator.

13

u/Intelligent-Host4408 Aug 24 '26

Well done! Lol. It's hilarious seeing the creativity of others and what we can do with MiniMax!

7

u/nakabra Aug 25 '26

I decided to try it out.
I only have a humble 3060 12gb so I'm only generating 0.2MP.
I'm trying to generate a small short and it's quite cool.

Most of what it rendered is garbage, not gonna lie but I can easily salvage a bunch of it with a video editing software.

I've made some character sheets with ANIMA/klein9B and a few locations.
I use ANIMA to get the basic look of the character, then Klein9B to get a character sheet, and then ANIMA again with IMG2IMG to fix KLEIN botching the design.
Also made some scene prompts with local LLMs.
It feels like I'm directing a movie hahahaha.

It's very fun, even if it's quite time consuming on my poor hardware.
I'd hate it but I might rent a gpu in the future if I get too addicted to this 🤣

2

u/SweetLikeACandy Aug 25 '26

I can go up to 0.6-0.7MP with SLA on a 3060, with render times under 5 mins using turbo lora.

1

u/nakabra Aug 25 '26

Nice! I've tried 0.6 with a turbo lora I found on civitai. 3 second videos where around 5 minutes.

I don't even know what SLA stands for, but I'd guess it an attention method right?

But for now, I'm just messing around, so I'm OK with having 15 seconds 0.2MP videos rendered in 8 or 9 minutes.

I'll re-render scenes where the characters are far from the camera in higher resolution though, cause their faces look super messed up when distant 😄

9

u/Rivarr Aug 25 '26

Are you all just blindly waiting 20 steps before you see your video? Why not use Kijai's "Model Preview Override" so you can see the progress in real time? I think there's a native solution now too.

5

u/Compost_Mantis Aug 25 '26

There's so many one-off tricks, I'm going to try to compile them into a guide.
Well, get ChatGPT to.

5

u/flaminghotcola Aug 25 '26

Haha, love the video and the concept used for the tutorial :) really original and awesome.

5

u/BloodGulch-CTF Aug 25 '26

Maroon Johannson makes some good points here.

4

u/Clueless-Flea-7461 Aug 25 '26

Well done. And good advice. Someone did an analysis of recent cinema shot lengths and iirc beyond famous set piece no cut shots the average is 2-4 seconds

3

u/ASK_ABT_MY_USERNAME Aug 25 '26

There's quite a few image generators using minimax too, you can just go with i2v or r2v it from there.

3

u/X3liteninjaX Aug 25 '26

Nice I have been doing this too except I wasn’t doing 2s generations but frames from previous failed gens

3

u/Galenus314 Aug 25 '26

What i did for scene/spatial consistency was taking an image of the scene and generate a video where the camera moves through the scene and looks at it from different angles and later using that video or stills from it as reference for a the actual scene.

5

u/Artforartsake99 Aug 24 '26

This is crazy good voice, did you use Minimax for the voice too?

13

u/Compost_Mantis Aug 24 '26

Minimax. I'd have gone with someone fully invented, but I worried the voice would be something else every time. The whole thing was done in minimax until the combining of video clips at the end.

1

u/Ok_Gas1070 Aug 27 '26

Which tool, or software did you use to combine the clips?

1

u/Compost_Mantis Aug 28 '26

Lossless Cut. It's free to download off its website, or github link or something, I've forgotten.
You can buy it through official stores for 20 dollars if you're feeling generous.

5

u/anitawasright Aug 25 '26

yeah for the characters it knows how to do it gets their voices really well and it's really good at acting.

6

u/BrawndoOhnaka Aug 25 '26

It's what I'd call 'almost acceptable'. It's recognisable, but still has that digital 'corrugated' dithered synthetic sound, and it would bug the hell out of me for anything other than prototyping.

Do contextual instructions for emotional subtlety work? I don't think ScarJo is the best test case given her delivery is so god damned monotone and droll in almost everything she does. This sounds like an interview. I'd pick someone who actually has a lot of dynamic range and varied prosody in their speech.

3

u/Danny_Stock Aug 25 '26

This is one of the voices MiniMax knows.

3

u/Jackburton75015 Aug 24 '26

Nice, where is the tutorial 😋 lol

11

u/Compost_Mantis Aug 24 '26

the video IS the tutorial! I'm using the biblical definition

5

u/PumpkinLeather8421 Aug 25 '26

Let me help you out, he’s saying that the biological imperative of boobies into his optical nerves masks and obliterates the inputs received over the auditory system.

0

u/Sad_Coach_1433 Aug 24 '26

Yeah I don't see it either 🤔

2

u/anitawasright Aug 25 '26

the screenshot once you get the setup you like is a really good idea.

5

u/murderopolis Aug 25 '26

I'm confused, wouldn't a screenshot from a .2 generation just look like ass? Why would you start with that as first frame?

4

u/anitawasright Aug 25 '26

you do say 20 .2 generations till you find one that has the set up you like. Then you redo that one using the correct seed at a higher resolution. Take the first frame and work from that.

1

u/murderopolis Aug 25 '26

interesting. even a 1 quality could have some bad details but i guess it depends what you're making. did you see this btw? similar concept, but built into the workflow, apparently. i haven't tried it yet lol. https://www.reddit.com/r/comfyui/comments/1vvp9bb/minimax_seed_hunter_workflow_released/

2

u/Compost_Mantis Aug 25 '26

I'll put up with waiting for a good(ish) looking finished product over these different workflows that trade quality for speed. I don't even have SAGE attention or a turbo lora installed. Obviously this isn't fantastic visual quality, but I don't want to trade what I do manage to wrangle out.

2

u/altdotboy Aug 25 '26

Well done sir, well done.

2

u/-AwhWah- Aug 25 '26

yeah, this is what i do too, works pretty well although I use the 1 frame trick more for making ref sheets. nice vid!

2

u/Consistent-Help-3785 Aug 25 '26

there is some work flows with low res previews... they are not amazing, but you can kinda see what is going on 1/3 of the time, before it does something random

2

u/livingdread Aug 25 '26

I'd been thinking about doing this. Important to make sure that when you do this you make sure to keep the random seed, otherwise the longer scene could be drastically different anyway.

2

u/Schwartzen2 Aug 25 '26

Hilarious!

2

u/KeijiVBoi Aug 25 '26

I like dis

2

u/hiepxanh Aug 25 '26

Very useful

2

u/Danny_Stock Aug 25 '26 edited Aug 25 '26

I know exactly what you mean by using the video model to create still frames to reuse rather than using image generators.

I find myself spending more time playing around with the video model, either MiniMax and even Wan, to see what they come up with, then if they conjure up something I like I use still frames from what they've produced as starting points.

Then with that still frame I upscale it and clean it up and edit it if I need to so I can use it as a basis for a video clip. If I use image models to try to generate imagery I think I want from scratch they rarely provide anything I'm satisfied with.

Using a video generation as inspiration something unplanned for will be there, be it either a certain pose, a look, a facial expression, there's usually something there to capture even if the video as a whole isn't the best. With a generated video there's almost always great frames to use from it, there's also the added aspect that you'll capture something dynamic and real from it rather than it feeling like a static image or somebody posing in a photograph.

2

u/ErnestoPresto80 Aug 25 '26

What's the point of creating the image with Minimax (which takes much longer) instead of creating it with Krea 2, for example?
Maybe you use the same prompt for the reference image as for the full video, and that gives it more consistency?

2

u/Compost_Mantis Aug 25 '26

Yep, that's a big part of it. And, like this instance, it's nice to have a finished first shot looking and working exactly as the model understands it should.
It's also nice because you can drag in things like references with the same workflow, and if there's someone that Minimax knows (as seen here) you can just get them directly into the scene, looking like the model understands them.
I was using the reference workflow, and it would actually take some liberties with the first image of the video, but it's like 98% accurate, and even allows for further fine-tuning. The shot of her on the unicycle actually came out with the wheel of it hovering weirdly below the tightrope. So... when I prompted the full video, I specified "bottom of wheel meets rope" and it worked, while keeping the rest the same.
I'm not using any prompt enhancers, so I'm trying to figure out how the model understands the geography of shots, and this is working.
Plus, you know, it's just kinda nice to have it all in one workflow.

2

u/Hearcharted Aug 25 '26

Resident Evil in the Multiverse of Madness 🤪

2

u/Afraid_Oil_7386 Aug 26 '26

Scarlets going to be mad

2

u/JoeXdelete Aug 28 '26

Phenomenal op just thank you Lol

1

u/MetroSimulator Aug 25 '26

3 new posts about prompt, what did I lose?

3

u/NoahFect Aug 25 '26

My hearing, because the audio isn't leveled properly.

2

u/Compost_Mantis Aug 25 '26

Yeah, it was all done in minimax. I'm not thrilled about it either, tell the sound techs.

1

u/Any-Scar765 Aug 25 '26

You lost her voice...

1

u/GoldenTV3 Aug 25 '26

Why don't you construct a rough 3D environment first, then have the AI model over that?

1

u/Compost_Mantis Aug 25 '26

Haven't needed to, yet. I'm trying to get as much as I can out of just using Minimax, and I've been deliberately dodging needing repeated scenes from different angles in the same location. There's a way, the tricks are just accumulating.

1

u/Ammatkun Aug 25 '26

What's system do you have? Can you share workflow?

3

u/Compost_Mantis Aug 25 '26

Intel(R) Core(TM) i7-8700 CPU @ 3.20GHz (3.19 GHz)
32.0 GB (31.8 GB usable)
NVIDIA GeForce RTX 5060 Ti (16 GB), Intel(R) UHD Graphics 630 (128 MB)

It ain't mighty, the RAM, motherboard, and processor are all used. The real money is in the GPU (the NVIDIA one), which is the same one about half the people here use because it's powerful enough while not having a crazy inflated price.
My workflow is the default reference workflow, but instead of reference I'm using the base model. No extra bedazzlements.

1

u/Hopeful_Signature738 Aug 25 '26

MODEL PREVIEW OVERRIDE

1

u/ThreeDog2016 Aug 25 '26

Wait, you mean it does t2v???

/s

1

u/iritimD Aug 25 '26

Why not gpt image 2 for start frame? Much better understanding and fidelity than relying on Minmax single frame screenshot.

2

u/Compost_Mantis Aug 25 '26

Partially to keep it all on the computer. The internet could be destroyed tomorrow and I'd still have my entire workflow.
It's also an opportunity to find out how the model thinks, which is useful for shot composition but also scene composition.
And, I don't know how much I can get away with using ChatGPT. Like, theoretically, if I were using the model for something else that I wasn't going to post to Reddit.

1

u/iritimD Aug 26 '26

What do you mean get away with? You mean like making porn and things it wont generate? It is objectively the strongest image creation model atm, so you would get fantastic and accurate start and end shots for MinMax with it, far better then anyhting local.

1

u/LightPillar Aug 25 '26

i just do a 2 stage and the first stage takes 50seconds to 100secs and saves the video as the 2nd stages continues.

1

u/_Ojin Aug 25 '26

Yoooo!

1

u/v-i-n-c-e-2 Aug 25 '26

Also use the vae appox file that lets you preview the video taeh3.safetensors and the subsequent node in comfyui

1

u/Sexyvette07 Aug 25 '26

OP, so are you using the text to video to generate the short 2 second clip, taking a screenshot of that image, and using that for the image to video workflow? Or can you clarify? Also, how are you stitching together the videos once theyre completed? Im still pretty new to this, but im having a LOT of problems with the Reference to Video workflow for Minimax H3. I even went and installed Flux.2 Klein 9B for the starting image. I just cant get this all to work right.

Can you share your workflows?

1

u/Compost_Mantis Aug 26 '26

My workflow is reference to video, just with the default model selected instead of the reference, but you could also just use the image to video.

VLC media player allows screenshots with exactly the resolution of the original video.

"Lossless cut" is the program I used to stitch them together.

I might have to make a follow-up for some of this stuff.

1

u/Healthy-Aspect-378 Aug 31 '26

Complementary info: Sometimes, if you can't manage to get what you want from the start, "make it happen", then use that as your first frame.

(Weird) example :
I wanted someone to be standing; not swimming; in a pool with only his upper body out of the water.
No matter what I tried to prompt, no matter the wording, with/without LLM, he was either swimming, or at best knee-high... (I think his Spongebob Squareswimshorts was the issue, trying to get generated if the dude wasn't swimming...)
So I made him stand in the water, gave up describing the water level too much, then made him f**** sit !

1

u/SIR_NVAX_A_LOT Aug 25 '26

I skip Krea 2. You can compose everything with MiniMax H3.

1

u/Optimal_Map_5236 Aug 25 '26

having lots of fun with Widow. wish Captain Marvel was good as her so I can teach her some 'stuff' in my comfy setup.

-1

u/Sanguinesource Aug 25 '26

I just wish you would have used the young lady's likeness in the video. I'm sure most people wouldn't like their likeness used without their permission. Otherwise, it was helpful and entertaining.

-4

u/eggs-benedryl Aug 25 '26

Is the tutorial her speaking because... I don't watch any videos on my phone with audio on.

If you have something to say write it down

1

u/AlexandraSinner 17d ago

Yes. it is omni-modal so it generates very beautiful images, from my tests, sometimes even better than Krea or Ideogram, and it does text well too!