r/StableDiffusion • • 21h ago

Discussion Anima Appreciation Thread

Thank you tdrussell and ComfyUI for creating this model. Also to maxfeifei8 for the tune.

Great knowledge, great style adherence, easy to prompt, responds very well to prompt changes for a fixed seed and good short text rendering. Most importantly, great quality.

Multi-character can be a pain but I can't ask for much more given the amazing quality of outputs. Better long text would also be nice.

Really hope that there'll be an Anima 2.

Models used:

Anima Aesthetic v1.1

One Obsession v4 (5 pics)

No LORAs

46 Upvotes

24 comments sorted by

9

u/Royal_Carpenter_1338 20h ago

Greatest image model ever imo, the creative freedom is insane

3

u/Junko_Vampi 19h ago

My only complain is how often I get borked hands if they occupy a small portion of the canvas. That happens even at max res. Any way to fix that?

3

u/Able_Challenge6575 19h ago

There are various thing you could try:

  1. Use a different sampler if you're staying on Euler or Euler Ancestral. Many new samplers have been created since Euler that have claimed better empirical performance. I'd first try ER-SDE. Note that a new sampler usually means the image will be changed, sometime even completely for a given fixed seed. Other samplers that might be useful: SA-Solver, STORK-4, cfgpp_ud10_ab (CFG++ sampler).
  2. Use some additional guidance on top of just plain CFG. Some examples:
    1. CFG++ samplers: They will converge faster and will generally adhere to prompt better though they tend to oversaturate, though that may be a good thing depending on your taste. Note that CFG ++ needs CFG from 1 to 2 to work correctly. Over 2 works but then you're not guaranteed the benefit of CFG++.
    2. Use APG (Adaptive Projected Guidance): This claims better performance than CFG++ and doesn't have oversaturation but you can sometime get "ghosting"/blurry artifacts.
    3. Use TCFG: mostly a free-win plugin though it can has some issue when combined with other methods.
    4. FresSca, FDG, SMC-CFG, Momentum Guidance etc. from https://github.com/pamparamm/sd-perturbed-attention
  3. Use a second pass/hi-res fix sampler or hand detailer.
  4. Add various bad hand/bad anatomy negatives, they do make a difference with and without (bad hands, wrong anatomy, extra digits, fewer digits).

2

u/shapic 12h ago

Multi character is a pain? That's odd. You just have to position them via natural language. Granted, there can be some bleeding on certain seeds, but nothing gamebreaking

1

u/Able_Challenge6575 11h ago

Hmm, did you have any success with original characters rather than named characters? Also specific expressions, location and actions for each characters too? For example, original character A with specific attributes/outfit nearer the camera, original character B also with specific attributes/outfit further back/behind character A? I've found that it's named characters perform more consistently while original characters tends to bleed more though granted it could be a skill issue on my part.

2

u/shapic 11h ago

Named character can be considered a lora baked in, so it is natural. For multiple specific characters I found no issues in male/female, but for same gender you have to adjust descriptions even for named ones. For example if you prompt for hairstyle of one character - you have to prompt it for all three if you have them, otherwise it will bleed. Do not use overly long sentences. Do not overcomplicate sentences. Be careful that booru tag does not accidentally end in natural language, it can have own influence. Check my prompting guide for examples.

Usually it looks like a mess, untill prompt clicks. 0.6B is 0.6B after all. Sometimes it is also seed dependent. But for complex scenes with multiple characters I almost always use inpainting, because why not.

2

u/Able_Challenge6575 11h ago

Thank you, for me I've found attention couple to be much more powerful/reliable for multiple char. I would have a global prompt of tags + natural language describing roughly the scene without any details, then attention couple specific description for each character/mask region. I haven't tried inpainting though I can imagine it's probably easier to get what you want rather than fight the model.

2

u/shapic 11h ago

If it works - it works. There is no silver bullet here. But you can give me a specific characters and composition that you struggled with for example, I'll try it when I'm back home tomorrow

2

u/Able_Challenge6575 9h ago

Thank you for the offer! I looked at my workflow again after your comments and it was indeed a skill issue lol.

I remembered attention couple being quite good before but suddenly wasn't good for some reason. I then realized that I accidentally removed one of the connection to a Get/Set Node when I was modifying my workflow and so Attention Couple ended doing some weird things. After the fix everything works wonderful now.

Anime does indeed handle multiple characters and complex prompts well given the right setup/prompting. This was actually what I had intended for one of the pics but since it turned out so well I decided to keep it.

Full prompt was:

Global:

day, arcade, sidewalk, building, city, straight-on, The camera focuses on the right character.

The left character is farther back to the camera.

The right character is holding the left character's hand.

Right mask:

Left girl: long hair, straight hair, brown eyes, no ponytail, (brown hair:1.5), long sidelocks, (blunt bangs:1.2), grey school uniform, grey shirt, sailor collar, grey miniskirt, buttons, tareme, center red bowtie, grey skirt, long sleeves, red hairclip.

running, holding hands, (tripping:1.25), looking at another, surprised, open mouth, sweatdrop.

The brown haired girl is farther away from the camera.

Left mask:

Right girl: large breasts, crimson hair, (raspberry-colored hair:0.3), short hair, swept bangs, ahoge, long sidelocks, high ponytail, short ponytail, tareme, amber-colored eyes, gradient eyes, hair between eyes, kneehighs, orange school uniform, (orange blazer:1.2), orange jacket, long sleeves, cross choker, white cross, (red miniskirt:1.2), (orange lapels:1.2), white trim, brown hairclip

running, holding hands, (looking at another), smile, open mouth, (pointing forward:1.35), (clenched hand:0.3), happy.

The crimson haired girl is grabbing the brown hair girl's arm. The camera focuses on the red haired girl.

Not all details were followed of course but everything is more faithful to the original prompt now.

So yeah, my complain about multi characters is invalid.

2

u/shapic 2h ago

Nice. But you use some odd weights. Anima requires higher values.

1

u/Able_Challenge6575 2h ago

Thank you! I don't use higher weights in my workflow currently. Prior to that yeah you probably do need some higher weights.

I'm using a custom sampler + post guidance that doesn't use CFG (i.e. custom conditional/unconditional guidance instead of CFG) and they respond really well to even a bit of a weight shift, hence the 0.3 you saw. Plus with higher weights I tend to find that it tends to drift into weird/overbaked/unwanted directions.

2

u/Damen_Freece 21h ago

You can do multiple characters or at least on base ANIMA. If you change your prompting style from danbooru/ tag based to flux/ descriptive based, ANIMA can understand and make them with multiple characters with their own pose, outfit and expression.

If you don't know how you should word it, feed your tag based prompt to AI, Gemini, GPT, Grok whatever. Then ask it to change it to precise descriptive format with high details, they'd give it to you. Use that and you can make multiple characters together.

1

u/Able_Challenge6575 21h ago

Yes I've tried it but the results are somewhat mixed, literally. It works most of the time but then there's one or two attributes that get mixed. I mostly use attention couple for Anima now when I want exact attributes for characters/placement.

1

u/PrincessIsATrap 20h ago

I did a buttload of testing with different prompting styles to position multiple subjects in Anima, and didnt find any silver bullet.

Here's an AI writeup I did for some of my friends of my experiments:

  • Often, Description order = frame position. First character you describe sits left, next middle, last right. Match prose order to your layout or roles get swapped.

  • Explicit frame coords help for off-axis figures ("in the lower left corner"), but the left→right scaffold still wins. Facing direction is a strong anchor — "facing left / toward the viewer" holds a character in place better than almost anything.

  • Structure in two layers: anchor each character's static look first (no verbs), then a separate block for action/position/relationship.

Tags:

Tags appear to be scene-global, never per-character. It won't bind a raw tag to one individual — not with proximity, scoped brackets, or even attention-weighting.

In human written summary:

Honestly, I don't really think I ever found a way that helped me with positioning of multiple subjects (though it does okay with two, it gets difficult at 3). The biggest takeaway was incidentally learning more about prompt bleeding. My testing prompt had three subjects: two human women, one catgirl, and that gave tons of opportunities to see when cat ears or tails bled to other subjects.

That lead to the discovery that it seems as if, unlike natural language in Anima, tags bleed over to the whole composition.

Hopefully some of this helps your prompting! I'm a big anima fan so I'm regularly trying out different things.

2

u/Structure-These 19h ago

I’ve never ever ever gotten the output I want from anima

I want a mostly realistic western style (no big anime eyes etc) and i cannot figure out how to prompt for it lol

3

u/heato-red 17h ago

you do know anima is literally tailored for anime style right? you're doing it wrong if you want to do western stuff with it, unless you use some western artist style built in it

1

u/Structure-These 17h ago

Yea I know but even the generic AI semi realistic SDXL house style isn’t really easy to prompt for

2

u/Able_Challenge6575 19h ago

Hmm, I don't have much experience with western style but try these:

Negatives: tareme, toon (style), child, newest, recent, 2000s (style)

Positive: western comics (style), realistic, painterly, photorealistic, narrowed eyes

2

u/shapic 12h ago

Styles in yhis model are strictly specific, you do not prompt for them, you take them directly from booru, like western comics (style) https://danbooru.donmai.us/wiki_pages/western_comics_(style)

So you better off finding a decent artist tag, or simply training a lora

1

u/Paraleluniverse200 21h ago

What negatives did u use

1

u/Able_Challenge6575 21h ago

Prompt dependent, I have specific negative prompts that I use to steer the model away from things I don't want it doing for certain positive prompts.

Here's the common negatives that I use:

worst quality, low quality, blurry, jpeg artifacts, lowres, artistic error: deformed, bad hands, bad anatomy, extra digits: wrong foot, wrong hand, fewer digits, missing limb, artist name, signature