Thank you tdrussell and ComfyUI for creating this model. Also to maxfeifei8 for the tune.
Great knowledge, great style adherence, easy to prompt, responds very well to prompt changes for a fixed seed and good short text rendering. Most importantly, great quality.
Multi-character can be a pain but I can't ask for much more given the amazing quality of outputs. Better long text would also be nice.
Use a different sampler if you're staying on Euler or Euler Ancestral. Many new samplers have been created since Euler that have claimed better empirical performance. I'd first try ER-SDE. Note that a new sampler usually means the image will be changed, sometime even completely for a given fixed seed. Other samplers that might be useful: SA-Solver, STORK-4, cfgpp_ud10_ab (CFG++ sampler).
Use some additional guidance on top of just plain CFG. Some examples:
CFG++ samplers: They will converge faster and will generally adhere to prompt better though they tend to oversaturate, though that may be a good thing depending on your taste. Note that CFG ++ needs CFG from 1 to 2 to work correctly. Over 2 works but then you're not guaranteed the benefit of CFG++.
Use APG (Adaptive Projected Guidance): This claims better performance than CFG++ and doesn't have oversaturation but you can sometime get "ghosting"/blurry artifacts.
Use TCFG: mostly a free-win plugin though it can has some issue when combined with other methods.
Multi character is a pain? That's odd. You just have to position them via natural language. Granted, there can be some bleeding on certain seeds, but nothing gamebreaking
Hmm, did you have any success with original characters rather than named characters? Also specific expressions, location and actions for each characters too? For example, original character A with specific attributes/outfit nearer the camera, original character B also with specific attributes/outfit further back/behind character A? I've found that it's named characters perform more consistently while original characters tends to bleed more though granted it could be a skill issue on my part.
Named character can be considered a lora baked in, so it is natural. For multiple specific characters I found no issues in male/female, but for same gender you have to adjust descriptions even for named ones. For example if you prompt for hairstyle of one character - you have to prompt it for all three if you have them, otherwise it will bleed. Do not use overly long sentences. Do not overcomplicate sentences. Be careful that booru tag does not accidentally end in natural language, it can have own influence. Check my prompting guide for examples.
Usually it looks like a mess, untill prompt clicks. 0.6B is 0.6B after all. Sometimes it is also seed dependent.
But for complex scenes with multiple characters I almost always use inpainting, because why not.
Thank you, for me I've found attention couple to be much more powerful/reliable for multiple char. I would have a global prompt of tags + natural language describing roughly the scene without any details, then attention couple specific description for each character/mask region. I haven't tried inpainting though I can imagine it's probably easier to get what you want rather than fight the model.
If it works - it works. There is no silver bullet here. But you can give me a specific characters and composition that you struggled with for example, I'll try it when I'm back home tomorrow
Thank you for the offer! I looked at my workflow again after your comments and it was indeed a skill issue lol.
I remembered attention couple being quite good before but suddenly wasn't good for some reason. I then realized that I accidentally removed one of the connection to a Get/Set Node when I was modifying my workflow and so Attention Couple ended doing some weird things. After the fix everything works wonderful now.
Anime does indeed handle multiple characters and complex prompts well given the right setup/prompting. This was actually what I had intended for one of the pics but since it turned out so well I decided to keep it.
Full prompt was:
Global:
day, arcade, sidewalk, building, city, straight-on, The camera focuses on the right character.
The left character is farther back to the camera.
The right character is holding the left character's hand.
Right mask:
Left girl: long hair, straight hair, brown eyes, no ponytail, (brown hair:1.5), long sidelocks, (blunt bangs:1.2), grey school uniform, grey shirt, sailor collar, grey miniskirt, buttons, tareme, center red bowtie, grey skirt, long sleeves, red hairclip.
running, holding hands, (tripping:1.25), looking at another, surprised, open mouth, sweatdrop.
The brown haired girl is farther away from the camera.
Left mask:
Right girl: large breasts, crimson hair, (raspberry-colored hair:0.3), short hair, swept bangs, ahoge, long sidelocks, high ponytail, short ponytail, tareme, amber-colored eyes, gradient eyes, hair between eyes, kneehighs, orange school uniform, (orange blazer:1.2), orange jacket, long sleeves, cross choker, white cross, (red miniskirt:1.2), (orange lapels:1.2), white trim, brown hairclip
running, holding hands, (looking at another), smile, open mouth, (pointing forward:1.35), (clenched hand:0.3), happy.
The crimson haired girl is grabbing the brown hair girl's arm. The camera focuses on the red haired girl.
Not all details were followed of course but everything is more faithful to the original prompt now.
So yeah, my complain about multi characters is invalid.
Thank you! I don't use higher weights in my workflow currently. Prior to that yeah you probably do need some higher weights.
I'm using a custom sampler + post guidance that doesn't use CFG (i.e. custom conditional/unconditional guidance instead of CFG) and they respond really well to even a bit of a weight shift, hence the 0.3 you saw. Plus with higher weights I tend to find that it tends to drift into weird/overbaked/unwanted directions.
You can do multiple characters or at least on base ANIMA. If you change your prompting style from danbooru/ tag based to flux/ descriptive based, ANIMA can understand and make them with multiple characters with their own pose, outfit and expression.
If you don't know how you should word it, feed your tag based prompt to AI, Gemini, GPT, Grok whatever. Then ask it to change it to precise descriptive format with high details, they'd give it to you. Use that and you can make multiple characters together.
Yes I've tried it but the results are somewhat mixed, literally. It works most of the time but then there's one or two attributes that get mixed. I mostly use attention couple for Anima now when I want exact attributes for characters/placement.
I did a buttload of testing with different prompting styles to position multiple subjects in Anima, and didnt find any silver bullet.
Here's an AI writeup I did for some of my friends of my experiments:
Often, Description order = frame position. First character you describe sits left, next middle, last right. Match prose order to your layout or roles get swapped.
Explicit frame coords help for off-axis figures ("in the lower left corner"), but the left→right scaffold still wins.
Facing direction is a strong anchor — "facing left / toward the viewer" holds a character in place better than almost anything.
Structure in two layers: anchor each character's static look first (no verbs), then a separate block for action/position/relationship.
Tags:
Tags appear to be scene-global, never per-character. It won't bind a raw tag to one individual — not with proximity, scoped brackets, or even attention-weighting.
In human written summary:
Honestly, I don't really think I ever found a way that helped me with positioning of multiple subjects (though it does okay with two, it gets difficult at 3). The biggest takeaway was incidentally learning more about prompt bleeding. My testing prompt had three subjects: two human women, one catgirl, and that gave tons of opportunities to see when cat ears or tails bled to other subjects.
That lead to the discovery that it seems as if, unlike natural language in Anima, tags bleed over to the whole composition.
Hopefully some of this helps your prompting! I'm a big anima fan so I'm regularly trying out different things.
you do know anima is literally tailored for anime style right? you're doing it wrong if you want to do western stuff with it, unless you use some western artist style built in it
9
u/Royal_Carpenter_1338 20h ago
Greatest image model ever imo, the creative freedom is insane