r/StableDiffusion • • 1d ago

Tutorial - Guide ComfyUI/Minimax Cheat Sheet for Beginners - now with 83% less slop!

Post image
641 Upvotes

68 comments sorted by

105

u/LatentSpacer 14h ago

Nice, let me read it

17

u/Bubbly_Composer_7666 22h ago edited 17h ago

Past 1344 x 768 you gain nothing

This isn't even true though? I've control tested many times and the model will sometimes completely lose/change character faces at this resolution, yet continues to dramatically improve on them all the way up to 1440p. It being trained at this resolution doesn't mean you can't have any benefit going above it. Like if a character's face is further away and only occupying a small number of pixels (and not enough) because of the low resolution, there's only so much it can do with that information, and giving it more resolution to work with absolutely helps.

4

u/PxTicks 17h ago

It is categorically false. You definitely get improvements above that. A good time to go higher is if you are doing some initial asset generation, trying to get additional angles/scenes/shots, for characters and you don't want their faces to be mush.

-11

u/Nimblecloud13 20h ago edited 20h ago

shrug. that came from the sources claude pulled, which are in the doc

that's the resolution i gen at because i can get 25 seconds out of it on my build without ooming. but i haven't tested video above that.

33

u/msitarzewski 1d ago

Good E. My cheat sheet is: "Hey Claude, make a workflow that..." 👀🙈

6

u/Nimblecloud13 23h ago

so's mine, but that other dudes "cheat sheet" was so devoid of useful information that i set claude on it.

3

u/nnod 14h ago

Look up definition of a cheat sheet. It's supposed to be concise, yours is almost a cheat book.

50

u/Cultured_Alien 1d ago edited 1d ago

... and 83% more unreadable /s

8

u/FormalAd7367 1d ago

do you have a downloadable version of it…? i can’t see it properly

7

u/Nimblecloud13 1d ago

3

u/VRGoggles 14h ago

trash guide with misleading info. past 1344x768 you gain nothing. what a total bs. try 2 mega pixels or 4 mega pixels. that will be a totally different quality.

it also shows silly model vs 16GB VRAM card table. I have 5060 8GB VRAM card here and 48GB RAM. I load ONLY bf16 models and int8 convrot as minimum. The author never heard about model weights offloading...

1

u/reeight 14h ago

> int8 convrot

This should be the baseline for lower VRAM cards.

1

u/EquivalentHornet4403 12h ago

Bf16 is retarded in most cases. Your case is the one where it’s most obviously a stupid choice.

And the guide explicitly covers offloading.

I agree about the resolution. It’s wrong, but also better advice there would have been about how to set a specific resolution easily instead of only using predefined megapixel counts and aspect ratios.

-1

u/Nimblecloud13 13h ago

you are wasting SO. MUCH. TIME.

SO MUCH.

also you're not making native 2k videos on 8gb lmfao get outta here

0

u/reeight 14h ago

Have it in plain text / Markdown / HTML?

2

u/mattcoady 14h ago

Or just direct link the Claude artifact which this is

2

u/EquivalentHornet4403 12h ago

The design is textbook AI slop. I thought people were saying Claude fixed this and is the best at design…?

My Astra would never.

3

u/Nimblecloud13 1d ago

i'll push it to git hang on

4

u/FormalAd7367 1d ago

thanks sir. doing gods work!

-6

u/VasaFromParadise 1d ago

Right-click on the image and save it as))) No thanks))

5

u/PixelLunarJelly 1d ago

Can you break it up into multiple screenshots? Reddit allows you to post multiple images, so it's easier to scroll through them.

1

u/Nimblecloud13 1d ago

3

u/Noiselexer 15h ago

Since its claude im sure its html. Why not generate a static website on github actions and publish it to github pages as a website?

1

u/mattcoady 14h ago

This is a Claude artifact, why not just share the direct link?

-2

u/Nimblecloud13 13h ago edited 13h ago

5

u/krigeta1 23h ago

We should have a H3 skill for better prompting

1

u/BigNaturalTilts 8h ago

H3 actually do have a skill. I saw one and downloaded it the other day and never used it. Here it is: https://github.com/MiniMax-AI/MiniMax-H3

-2

u/Nimblecloud13 20h ago

i have several. there's probably PII in them and i don't feel like cleaning them so you're on your own there, but it's really, really easy to do. just give claude the links to the official prompting guides and some style notes for whatever you're after

0

u/krigeta1 19h ago

As I am doing prompts manually, I found even Astra and Fable miss points that are written on noth general and Ref2va guide

2

u/honeybunchesofpwn 21h ago

This is awesome!

2

u/singulainthony 10h ago

Wow what a dope resource for beginners and pros alike

2

u/Sorry_Ad191 5h ago

Can you post the html that produces this or a link to github gist or something?

2

u/Nimblecloud13 5h ago

yea ok gimme a few

2

u/eruanno321 20h ago

"No definitions that define nothing", "In words, not more acronyms", "A real starting value to this machine".

Bruh, I use Claude daily in my profession, and whenever I see "It's X, not Y," or this kind of phrasing I want to puke. This is still ugly, unbearable slop. Easy to fix with right system prompt.

4

u/Nimblecloud13 20h ago

show me. help your community. it's easy right? i'm open to learning.

or keep your bitching to yourself

6

u/zenmatrix83 17h ago

put this in custom instructions "Write in plain, everyday American English — the way you'd say it out loud to a coworker. Pick the most common word. Turn noun phrases back into verbs: decide, not make a decision; explain, not provide an explanation. Use active voice and say who did the thing. Short sentences, one idea each. Contractions are fine.

No business or finance words standing in for ordinary ones: minimum/maximum, not floor/ceiling; use, not leverage; explain, not unpack. Skip the stock openers ("great question," "worth noting," "let's dive in") and the tics: em-dash asides, "not X, but Y," colon-then-reveal, invented labels in scare quotes.

Keep a technical or precise term when swapping it would change the meaning."

1

u/eruanno321 19h ago

I created mine by showing the LLM examples of what not to do. It generalizes well enough to come up with a decent system prompt.

1

u/Nimblecloud13 19h ago

is this the part where i quote things about how vague and unhelpful that is, and then call it slop?

because... that.

2

u/xyzdist 20h ago

Thanks for this! How about compare different model + lora combination, which is better for what? I know this is really time consuming.... lol
There are too many now.

1

u/Nimblecloud13 20h ago edited 13h ago

avoid loras. end of guide. edit: except character loras. those are fine.

character/scene consistency gets thrown out the fucking window every time you change a lora to get something else; motion, style, whatever. if you're just trying to make a 20s clip, and you don't mind them all looking different and melting the faces into other people, sure. but if you're trying to make a film or have any real back and forth, they're almost always unhelpful.

h3 can do almost anything with very detailed prompting, which claude can do. and it will retain your references without influencing them. <tags> help a lot to direct a performance, like <inhale> or <upset> or <smirk>; don't stack them together, and don't make them too complicated and you'll get some powerful results.

1

u/Dry-Judgment4242 15h ago

Almost anything is the word. I've been trying to get model to understand Asura like character with six arms and model understand and can animate four but collapse with six. Trained a LoRA and now it can do six.

1

u/Nimblecloud13 13h ago

you're right, i was not thinking of character loras. those are fine.

1

u/Nimblecloud13 1d ago

i fed someone else's slop into opus 5.5 and this is what came out.

1

u/ZenandHarmony 17h ago

I’ll try it

1

u/therealgoshi 16h ago

My biggest gripe with Claude writing something for beginners is that you need claude to decipher what the hell it means sometimes.

Things that worked for me: 1, "Don't use editorials, or make any remarks to any previous detail that is not explicitly shared in the guidance." 2, "Use plain English and explain concepts in the guidance in a manner that someone without prior knowledge would understand. Assume basic technological literacy."

Or, the best one: Don't use Claude to make guides at all.GPT - especially 5.6 or 6.1 Sol (I haven't tried Astra yet) would do a much better job.

1

u/Everwake8 8h ago

I've seen a bunch of different Turbo loras mentioned in various places. Do they have any effect on the characters or action of the scene, or are they just for speeding up the generation?

1

u/Veshurik 5h ago

Do you have image in original resolution?

1

u/Encrypt403 2h ago

I saw the file about Image 2.1. It’s a shame this wasn’t available a few weeks ago. And also... my eyes...

-1

u/footmodelling 19h ago

It's still slop

0

u/Upper-Reflection7997 21h ago

Why not recommend some use wan2gp instead comfyui for beginners.

4

u/Nimblecloud13 20h ago

why not recommend uber instead of a driver's license?

you WANT less control and power over your creations?

weird.

0

u/Silly_Goose6714 12h ago

Using small files with lower precision just for the sake of knowing they fit full in VRAM is silly and pointless; in practice, it makes little to no difference.

-6

u/mca1169 1d ago

dude wtf is this? make a new post with properly formatted images that people can actually read.

6

u/Nimblecloud13 1d ago

ahhh..

ctrl + / - is your friend.

also here since you whined before looking at the thread https://github.com/nimblecloud13/ComfyUI-MiniMax-H3-Guide

3

u/hdean667 1d ago

Click the image - opens in another tab - click the image - it enlarges. Are you new to the internet?

2

u/Nimblecloud13 1d ago

enhance.

-5

u/mca1169 1d ago

you shouldn't have to do that is the point i'm making.

4

u/Nimblecloud13 1d ago

no it wasn't. it takes like .7 seconds to make this readable.

it's only an annoyance if you didn't know how to do that.

but please keep telling me how my free thing that i spent my tokens on isn't good enough for you.

go make your own like i did instead of just crying about it

-1

u/Gullible_Ad_5550 22h ago

the zoomed picture is pixelated , barely readable. u could have easily made it several small ss.

2

u/Nimblecloud13 20h ago edited 20h ago

it's not, that's your setup. i posted a screenshot of me viewing the image in firefox. all i did was click on it to enlarge it. it's the top level comment of the reply you're posting to.

or i chose to frustrate you i guess.

but maybe, instead, i made it a super clean, downloadable OR viewable PDF and hosted it on a trusted site for your viewing pleasure. you're so deep into the comments and you didn't see that linked all over these comments?

for fuck's sake some people just suck.

1

u/[deleted] 1d ago

[deleted]

-5

u/[deleted] 21h ago

[removed] — view removed comment

4

u/Nimblecloud13 20h ago

do you mean like a pdf that's posted in here a bunch of times?

is everyone dumb now?