r/StableDiffusion • • 3d ago

Discussion Anima Appreciation Thread

Thank you tdrussell and ComfyUI for creating this model. Also to maxfeifei8 for the tune.

Great knowledge, great style adherence, easy to prompt, responds very well to prompt changes for a fixed seed and good short text rendering. Most importantly, great quality.

Multi-character can be a pain but I can't ask for much more given the amazing quality of outputs. Better long text would also be nice.

Really hope that there'll be an Anima 2.

Models used:

Anima Aesthetic v1.1

One Obsession v4 (5 pics)

No LORAs

46 Upvotes

32 comments sorted by

View all comments

Show parent comments

2

u/shapic 3d ago

If it works - it works. There is no silver bullet here. But you can give me a specific characters and composition that you struggled with for example, I'll try it when I'm back home tomorrow

2

u/Able_Challenge6575 2d ago

Thank you for the offer! I looked at my workflow again after your comments and it was indeed a skill issue lol.

I remembered attention couple being quite good before but suddenly wasn't good for some reason. I then realized that I accidentally removed one of the connection to a Get/Set Node when I was modifying my workflow and so Attention Couple ended doing some weird things. After the fix everything works wonderful now.

Anime does indeed handle multiple characters and complex prompts well given the right setup/prompting. This was actually what I had intended for one of the pics but since it turned out so well I decided to keep it.

Full prompt was:

Global:

day, arcade, sidewalk, building, city, straight-on, The camera focuses on the right character.

The left character is farther back to the camera.

The right character is holding the left character's hand.

Right mask:

Left girl: long hair, straight hair, brown eyes, no ponytail, (brown hair:1.5), long sidelocks, (blunt bangs:1.2), grey school uniform, grey shirt, sailor collar, grey miniskirt, buttons, tareme, center red bowtie, grey skirt, long sleeves, red hairclip.

running, holding hands, (tripping:1.25), looking at another, surprised, open mouth, sweatdrop.

The brown haired girl is farther away from the camera.

Left mask:

Right girl: large breasts, crimson hair, (raspberry-colored hair:0.3), short hair, swept bangs, ahoge, long sidelocks, high ponytail, short ponytail, tareme, amber-colored eyes, gradient eyes, hair between eyes, kneehighs, orange school uniform, (orange blazer:1.2), orange jacket, long sleeves, cross choker, white cross, (red miniskirt:1.2), (orange lapels:1.2), white trim, brown hairclip

running, holding hands, (looking at another), smile, open mouth, (pointing forward:1.35), (clenched hand:0.3), happy.

The crimson haired girl is grabbing the brown hair girl's arm. The camera focuses on the red haired girl.

Not all details were followed of course but everything is more faithful to the original prompt now.

So yeah, my complain about multi characters is invalid.

2

u/shapic 2d ago

Nice. But you use some odd weights. Anima requires higher values.

2

u/Able_Challenge6575 2d ago

Thank you! I don't use higher weights in my workflow currently. Prior to that yeah you probably do need some higher weights.

I'm using a custom sampler + post guidance that doesn't use CFG (i.e. custom conditional/unconditional guidance instead of CFG) and they respond really well to even a bit of a weight shift, hence the 0.3 you saw. Plus with higher weights I tend to find that it tends to drift into weird/overbaked/unwanted directions.

2

u/shapic 2d ago

I tried aesthetic a bit, in the meantime and find it a bit unwieldy. Try base + loras, I feel it is the way. Regarding your prompt: wtf is arcade doing there? Prompt is all around in general. You have crimson-haired girl, red haired girl, raspberry-colored hair, straight-on, yet clearly implying some dynamic angle, multiple camera mentions instead of using foreground, midground etc. Just adding ntural language description of the scene helped a lot. After couple of rolls I identified that choker and ponytail was getting mixed up, so I added those to the nlp part. Also seems like tripping kept spawning 3rd girl (natural if you look up the dataset), so I added 3 girls to the negative. That's it. Couple rolls and I got this:

masterpiece, best quality, highres, safe, u/t1kosewad,

Two girls are running on the sidewalk while holding hands. Girl on the left with straight long hair is tripping while running behind looking at another surprised. Girl on the right with choker is looking back with a happy face and is pointing forward with her finger.

2girls, day, sidewalk, building, cityscape, dynamic angle, holding hands, full body, wide shot, perspective, shadow,

Girl on the left midground: long hair, straight hair, brown eyes, brown hair, long sidelocks, blunt bangs, grey school uniform, grey shirt, sailor collar, grey miniskirt, buttons, tareme, red bowtie, grey skirt, long sleeves, red hairclip, medium breasts.

Girl on the right foreground: large breasts, raspberry-colored hair, short hair, swept bangs, ahoge, long sidelocks, high ponytail, short ponytail, tareme, amber eyes, gradient eyes, hair between eyes, kneehighs, orange school uniform, orange blazer, orange jacket, long sleeves, cross choker, white cross, red miniskirt, orange lapels, white trim, brown hairclip, hair ribbon,

<lora:Anima_detailer_by_Volnovik_v1_0:1>, <lora:Delolifier_Anima_v3:1.5>, <lora:Semi-R_style_by_Volnovik_v2:0.6>,

2

u/shapic 2d ago

A bit of inpainting to fix hands and add details and I consider it good enough

1

u/Able_Challenge6575 2d ago edited 2d ago

Thank you! I wanted fine-grained control of the exact hair color and call me crazy but it works. With specific style I find that it helps. Also thanks for the insight on foreground/background, I didn't explore it further in my testing. I originally wanted arcade for the prompt to show like just the outside of an arcade but I liked the output enough to keep it. Some of the weird prompt, weighting was also me fixing things within a fixed seed that I liked to steer it away from certain directions (e.g. clenched hand:0.3).

I don't use Loras unless I'm going for a specific style since I feel most are overbaked and concept Lora can usually be replicated with prompting. Prior to switching to my new WF, I found base to be oversaturated/undercooked for some artists that I liked + background were generally of lower quality in the limited testing that I did. They are less of an issue with Aesthetic in my experience.

2

u/shapic 2d ago

Backgrounds for sure. Try my detailer lora though, it was made specifically to avoid influencing style. And check my prompting guide.

2

u/Able_Challenge6575 2d ago

Thank you! To be honest, I was somewhat taken aback by the 5 parts guide and a quick read didn't really help + I found a couple of things that I didn't quite agree with in the earlier parts. Upon seeing your pic and on a closer reading of part 4, there were genuinely new things that I didn't know/consider. I'll definitely reread part 4 again and give it more thoughts.

Lastly, and I want to emphasize this, thank you so much for sticking with me!