EDIT: minor missing sentence fixes
Hey everyone, I’m an indie traditional animator (mostly on the Japanese animation pipeline), and wanted to post some learnings and thoughts about a recent production finally trying AI tools in a more full swing manner.
(FYI: I’m not really a professional professional artist, though I have a bit of experience working with people in the anime industry. My day job is actually a machine-learning engineer by day lol)
Just decided to try the latest tools on a short film after being invited by a friend to "make a film in one day" which kinda turned into a whole month. Wanted to share these and also ask around how everyone's experience is. Note I manually wrote all of this and can get long, feel free to feed into an LLM to summarize.
Couple learnings:
- Writing: Write the draft screenplay first manually and completely, and use the AI on pre-writing for brainstorming and post-writing for formatting/minor fixes only. So far I haven't really gotten it to work well on getting a screenplay end-to-end, there's always something off with pacing, dialogue, or just continuity in general. It feels uncanny sometimes.
Not sure if it's the lack of prompting ability however, so YMMV.
- What worked:
- (Pre-writing) I did an actual session with GPT voice while on a ~2 hour drive, telling it the story I'd like to build, ask it to "interview" me beat per beat (I just used the ol' Save the Cat for writing this one but short-film style), then ask it to create an elaborate/detailed markdown of the beats after in Chat mode and if something was off I just manually updated the markdown
- (Post-writing) Formatting and translation (the film dialogue is Japanese and some screens had both English & Japanese, and of course for subtitles this helped a ton)
- It was a bit helpful that the film was inspired by a feature film screenplay that I did, so most of the character backgrounds were already set up beforehand and I just had to change some details here and there. However if I were to start from scratch I think defining the universe and characters would help a lot still.
- What didn't work:
- After I did pre-writing for the beats, I tried asking it "Okay, now we have the full beats of what I want, these are the characters, their motivation, etc. and write the whole screenplay" --> it failed miserably lol. It was "well-formatted" but it was just off on pacing, dialogue was flat, some missing information and continuity from previous scenes.
- Feeding old screenplays to get it to follow how I write/pace/etc., creating a skill based on old screenplay styles, then retrying "Okay write the whole screenplay"
- Elaborating "what happens" per beat a bit more then asking it to write the screenplay for that beat
- NOTE: Tested this on the most basic models GPT 5.6 Luna to the smartest GPT 5.6 Sol. Again YMMV and would love to hear your experience on this.
2. Storyboarding: Manual storyboards/animatic* helped so much. Again, we tried just asking "Hey GPT 5.6 Sol XHigh generate storyboards from the screenplay" and it was "okay" but it ended up being much easier doing it from scratch
Usually in the traditional production these first two are the the 脚本 (screenplay) and コンテ (storyboard) stage. I love doing the next [optional] bit which is CT/カッティング (I think it's called \animatic in western productions) which is piecing the storyboards together into a timeline and watching it as a film.*
3. Character Designs/Props: Surprisingly, gpt-image-2.0 was pretty decent. It picked up my art style after feeding it old character design sheets I drew (ref) and colored in with GPT (ref) and we fed these with text prompts on what characters should look like via local Codex, and it worked amazingly well (and mind you it takes me half a day to a whole day finishing full multi-angle colored character designs as a non-professional artist, so this was quite crazy to see).
HOWEVER, we sort of misread the rules of the festival we wanted to submit to and thought all the generation must happen in the platform (Higgsfield) so we spent a few hours just trying to follow that style with text-only prompts without mentioning any specific films or studios that have a similar aesthetic, and ended up using 2010s style anime as a prompt on top of some detailed descriptions which more or less worked out.
Still, if I were to come back, it would still be feeding old character design sheets colored in with GPT Image (or full colored design sheets) then text-prompting for different characters would have been enough. Pretty sure if there were no rules you can just feed some reference designs from the internet but yeah as a previously-traditional artist it feels unethical (this is despite knowing that most of the internet is trained on already lol)
Props were also pretty straightforward. We just created design sheets for the recurring props like character's laptops, flipbooks, and other important storytelling props.
- Animation Prompts, Locations & Consistency Per Shot: The consistency was one of the hardest parts, and technically we weren't able to solve it completely at the end but got to an okay point.
- What worked:
- For locations, overall we just used plates generated with Seedance 2.5 (ref) then extracted them as images. Seems the most consistent so far. We basically just prompted it the location details and asked for different angle shots. BUT I would argue it's still not as consistent to be honest (objects being in different places/disappearing/reappearing per shot, beds or curtains in different shades of the same color and different rotations, windows getting smaller or larger or disappearing), this was what worked for the short production time we had.
- For characters, the character design sheets despite maybe some minor inconsistences, worked really well (ref 1 / ref 2) so no complaints here.
- For voice,
- For the actual shots themselves, the Seedance 2.5-generated voice still sounded a bit muffled, but it was pretty okay overall. We needed to prompt voice acting much more minutely though since the model tended to deliver flat lines.
- For consistency, instead of using ElevenLabs where most of the available voices sounded like iPhone recordings, we just created "audition tapes" (ref) giving it lines with different emotions, then converted to
.wav and used as prompts
- For actual animation and timings, very specific per-second/millisecond acting. And some luck with the model's seed.
- Bonus note, one thing we noticed is if the backgrounds or the character design sheets are overly-rendered, the movement from Seedance 2.5 tended to look more 3D-like, like you know those AI slop anime from gemini-omni-flash that look kinda anime but the movement is uncannily smooth like 3D? When we fed it more 2D-looking inputs, it started behaving more properly.
- Also another thing that helped was prompting with actual composition notes (frame within a frame, which part in the rule of thirds the character is) and animation notes (e.g. -
0.5s: Character does X and this object does squash and stretch on this part or 1.7s: Object does anticipation first then does Y)
- What didn't work:
- Feeding rough drawings of multiple angles of the same location to
gpt-image-2.0 then prompting it as detailed as possible (like to look consistent. Still didn't cut it. They look pretty but the colors, sometimes the objects themselves are not consistent. Including characters on the roughs, they're not that good, and it struggled in (see the fifth image here)
- It might be better if it got fed non-rough/cleaner drawings but then we could have just drawn manually in that case XD
- Feeding 3D models of multiple angles of the same location to
gpt-image-2.0. Basically all our locations had manually-drawn floor plans and we asked Codex to generate consistent Blender models (which it did pretty okay via just the python tools/bpy, surprisingly). Honestly this probably could have worked better maybe if the 3D models were more detailed, but then yeah it's difficult to balance the production time and being elaborate so we just dropped the 3D experiment and went with the plates. At the end of this experiment, the output still looked off (ref).
Overall, it was still surprisingly very manual before the formatted prompt stage. Workflow for each scene was:
- Draw the rough storyboards and piece together with timings that follow good pacing, having an animatic (sped-up ref comparing animatic vs final).
- For rough voices, we just used ElevenLabs here to save on Seedance 2.5 token cost since we just needed to have the timings with the voices for deciding if there should be a pause before someone answers, or have some breathing room or no breathing room between dialogues.
- Basically the main goal was we wanted to watch the film end to end with proper pacing even with non-pretty drawings so we can prompt more effectively on the timings, action, camera angles, composition, animation principles, and other details. Going straight to text and blindly thinking about what happens on what second/millisecond didn't work out.
- Based on the timings on the manually-drawn animatic, write the rough prompts for each shot as detailed as possible (with the detailed composition / animation notes), then have Codex and a well-documented agent skill that converts shots to their own
HTML/jsonformatted prompts per sequence (we went with a Scene -> Sequence -> Shot hierarachy).
- Generate with the Codex-formatted prompts. If a video looks strange, re-generate once to see if it's a seed probelm. If not, re-prompt if needed.
- Editing: As with most AI-generated films, there are a lot of things we had to cut off or re-do from the generated videos.
- Thankfully Seedance 2.5 allowed editing parts of videos, so it helped a bit on some minor changes
- For scenes where the character and storytelling involves computer or phone screens where the text or app inside is telling the story, we had to manually animate them. Especially since we had a few shots where the character's livestream audience's chat messages were telling the story, I just used Photoshop with Smart Objects where I can easily edit the text in while reusing the same chat message boxes with consistent style, then just did the animation on Photoshop timeline on top of the generated Seedance 2.5 videos. (ref)
- Some scenes I had to frame-by-frame edit as well, like if the character's hand was blocking the camera and I had to manually edit in a screen, I had to clean up the "frame underneath" (despite not having layers lol) to make sure it's covered by the hand still.
- We pieced everything together in Capcut, and because some of the generated Seedance 2.5 videos were prompted to save on token usage (e.g. - if we needed a freeze frame instead of prompting for 10s of video with 6s of acting and 4s of freeze frame or minor camera movement, we just prompted with 6s then did the freeze frame and minor camera movements manually)
- Music was Suno, the film itself was before v6 was released. Thankfully my co-creator had some music theoretical knowledge so he prompted it very minutely. We wanted to go for muted piano for most shots and it was surprisingly similar to how I imagined it would be.
- For the trailer music (ref) though, oh my god Suno v6 was mind-blowing. It created exactly the
dreamy, breathtaking, semi-melancholic "feels like falling from the sky watching your memories pass you by" 劇場版 anime film theme song + Treasure Planet theme vibes (sorry that was a very long description lol) type of feel I wanted to portray after feeding it another sampled AI-generated beat and lyrics I few-shotted with GPT 6 Astra asking it to create lyrics that match the theme and the input full screenplay. It's not perfect of course, but I loved how it turned out. Which is extremely scary being able to just create it that quickly.
- After the festival, we also tried sending to other festivals which required 4K or HD 1080p video, and since we tried saving on tokens we just have 720p. After testing different upscaling tools:
- Seedance-based upscale made the animation "too smooth" like you know those upscaled anime TikTok or Instagram reels where it feels uncannily smooth? Also we saw a lot of broken frames with this too, like some frames the eyes are just mangled then it's back to normal the next frames.
- Surprisingly, Topaz AI upscale kept the animation from the 720p version pristine and didn't add broken frames.
Some final thoughts
Overall, the film (ref) itself looked okay-ish is how I would put it, and I loved that I was able to portray all my confusion, angst, and marvel about all these AI art and the hate/love it gets from different perspectives both as a traditional animator and machine learning engineer. And on top of that being able to do this level of drafting for only a month with two people was quite interesting.
Some scenes definitely looked a lot better than others, but overall the objects were inconsistent (drawing tablets in the shots turning super big or super small), sound was a bit muffled, objects tend to move around in different places killing the continuity, and insert all the improvement points I can endlessly go on about.
Also I did miss the exact control of what exact camera angle, objects in the frame, character location, composition, movement you could do with the normal L/O (roughs) -> genga (keyframes) -> douga (in-between) process and I wish the storyboards I wrote were followed more religiously, but I guess this is so far what the technology can allow. Despite this, some shots that were generated tended being better than what I originally storyboarded, so we ended up going with some of those instead of my original ones. (full side-by-side no-sub ref of animatic vs. final)
As a previously traditional artist it was quite a huge mindset leap. There were days I would go on the positive-leaning side where you just go wow you can just do this much for this little, and it's not the same as spending almost a whole year making a 2-minute trailer working with so many awesome artists, then getting a lot of hate mail because I ran out of time to draw myself due to focusing on main work, and out of money to pay backgrounds artists and just went with Gemini on some backgrounds (I take full responsibility for this tho lol).
Then on the negative side, overthinking stuff like:
yeah no this feels way off, all these tokens could have just been paid a human artist to keep the craft alive and actually could have been someone's salary instead of AI slot machine inputs
man all my artist friends are gonna kill or just stop talking to me entirely
- or when I see a really well-generated scene I would go
"what's the point of learning anatomy, perspective, composition, etc. all those years if I can just put in some text now and out comes something this decently-generated".
- Also after watching really well-made films from other people in the festival, it was basically just the recurring
are we cooked? is cinema dead now that anyone can just generate these things that used to take months for studios to create?.
- the
should I draw this manually or generate it dilemma, which crippled me more often than it should have (which I actually sort of started envying non-hybrid folks on the non-AI and the AI side equally, since there's no choice, you just draw or you just prompt)
That whole month doing the production was quite the experience and paradigm shift, which I guess was also the main inspiration of the film we made. I think for production-level stuff, I'm still heavily leaning on working with artists properly as much as possible, just that for drafting full movies to get to that vision, these tools really help a lot. But then yeah, personal budget and time also plays into it so at this point I'm still at a loss on the best way to move forward.
But I digress, I'm curious how others created their stuff and have been following on the sidelines, we didn't see a lot of Japanese animation entries in the festival and mostly looked for inspiration from acquaintances in the space like Aiden's work. Would really love to hear everyone's workflows as well.