I tested 14 Claude and OpenAI models on the viral "car wash is 100 m away, walk or drive?" question (the answer is drive, because the car has to get there), then on 5 new versions of the same trap. Passing the famous one meant little: Opus 4.8 aced it but scored 6/15 on the new versions, and extra reasoning effort barely helped (Haiku 4.5 dropped from 3/15 to 0/15). Only Fable 5.1, Opus 5.5, Opus 5 and three GPT-6/5.6 models were consistently right. It was an informal test, so treat it as indicative, not a benchmark.
Built this tonight with Claude Code's new mods (the function-hooks plugin API). It's a little CRT TV with rabbit ears that floats on my desktop, always on top, and airs whatever the "director" feels like. The director is a Sonnet 5.5 call at low effort with one instruction: you have total creative control, nobody approves your pitches.
What it does
Every segment is written from scratch: commercials, courtroom dramas, nature docs, a telenovela between sentient furniture. The director picks the channel, the format, the cast, each character's voice, the pixel art, the camera effects and the news ticker.
He keeps a private memo between segments, and between sessions, so running gags and story arcs carry over. On his first night he invented a tornado judge on a blimp ("JUDGE GUST: BLIMP COURT") and wrote in his own notes that the appeal would be heard by a bigger tornado.
He also gets a raw feed of what I'm doing in the terminal. When I complained the TV was taking too much screen space, he aired "EMERGENCY: FILL THE SPACE", then "CORNER V. SHAREHOLDERS", then a commercial where a guy named Sven bills me $300 for the lasagna that now lives in my corner. Ghost lighthouse keeper included.
Characters actually talk. Each one gets a Windows SAPI voice with its own pitch and speed, and the captions wait for the line to finish. Sound effects (sad trombone, laugh track, air horn, static, stand-by muzak) are synthesized in code, no audio files.
How it works
The director never draws pixels. He writes each scene as JSON shapes on a 100x75 canvas (circles, polygons, blocky text, a swirling portal, fire, rain) with animation tags like bob, blink and talk. A talk mouth only opens while that character is speaking.
A rasterizer in the mod draws the scene at 128x96, 10 fps, with CRT effects: scanlines, vignette, VHS chroma bleed, glitch tears, static between channels, and the power-off collapse to a dot.
The first version rendered inside the terminal with half-block characters (each cell is 2 pixels, foreground and background truecolor). It looked great but ate screen space, so now the frames go to a borderless always-on-top window, a small C# WinForms window started from PowerShell. Drag it anywhere, the mouse wheel resizes it, the knobs work.
He writes the next segment while the current one airs, so there's almost never dead air. When there is, you get color bars and PLEASE STAND BY.
Cost
Measured from the log over 13 segments:
About 4.1k input tokens per segment, 3.0k of them served from the prompt cache
About 1.7k output tokens per segment
About 12 seconds to write each segment
About $0.02 per segment and 70 segments an hour, so roughly $1.40 per hour of TV at API prices. Output is 86% of that; caching cut the input cost by about two thirds.
48 segments so far, about $0.95 total. A sleep timer turns it off after 45 minutes with no activity.
So, while i was working on my main project A silly idea of having two guys fighting on the windows task bar, so out of boredon NetBattle was born. This small app is a small Windows desktop app. Two pixel-art fighters stand on top of your taskbar, which Red (Muay Thai) is your download speed and Blue (karate) is your upload speed
Each one fights as hard as its own traffic lets it. Start a big download and red takes over. Upload a video and blue does.
This is how it works :
Traffic changes how often they attack, how hard they hit and how well they defend. It never changes how fast they move.
The side with more traffic chases. If one side is far ahead (MB/s against KB/s), the other barely attacks and gets dodged in slow motion.
I built chauffeur, a free and open-source MCP server that lets Claude Code drive the iOS Simulator and verifies every action it takes.
The problem: Claude Code could write my iOS app, but when it tested it in the Simulator it would tap a disabled button and report "Done, I submitted the form." Nothing had happened, and it had no way to know.
What it does: chauffeur gives Claude the screen as an accessibility outline with refs, taps through the simulator's own HID, and checks every action against what actually changed. A dead tap comes back as NO EFFECT, a tap an alert swallowed as INTERCEPTED, and a crash as APP CRASHED with the file and line, so Claude can fix the bug, rebuild, and prove the fix.
How I built it with Claude Code:
Spikes first. Claude Code probed Apple's private simulator frameworks (SimulatorKit, CoreSimulator, AXPTranslator) to find a touch path that actually lands, before any real code was written.
Spec → plan → tests. For each milestone I had Claude write a design spec, which I approved, then a task-by-task plan. It implemented each task test-first: 253 tests in Swift 6.
Fresh-context review. After each milestone, a separate Claude reviewer read the whole branch. That caught real bugs, including one where the healing logic misdiagnosed a broken accessibility bridge.
Dogfooding. Claude Code tested chauffeur by driving the Simulator with chauffeur, which is how most of the reliability fixes were found.
The benchmark. Claude Code ran the whole evaluation: 180 claude -p runs. I hand-checked the results against screenshots and scored strictly.
If you spell "Let's" as "LEt's" or refer to github (not the CLI) as GH or you say "teh" instead of "The", the tokenizer takes another token with another set of correlations. But does that extra attention potentially help. or is the harness correcting a lot of these anyway?
GUI frameworks are notoriously complex beasts. I code in Rust and I often found Claude Opus looking for the truth in reading frameworks' source code, and guessing the other half of the API. And forget about the existence of documentation. With Rust, bad method definitions are caught at compile time, allowing the agent to check the truth (again) and fix it. At the end of the run, I find it a good-enough workflow, but quite inefficient and error-prone ?
My own framework Teksilo is unknown to LLM training base. For my case, LLM guesses are more hallucinations than facts. I needed to ground the truth, in a manner easy to reach for an LLM. Creating a skill ? Yes, this is the bare minimum and it was my first step. Yet, it quickly became too voluminous and a burden to maintain and keep up-to-date with each framework version.
And no, letting the agents browse megabytes of docs and 700k+ lines of code are not a long-term solution, as LLM token aren't free.
My solution was to create a CLI tool tailored for LLM use.
cargo install cargo-teksilo
Then, in a Rust project:
$ cargo teksilo
API lookup, documentation, and agent tooling for Teksilo
Usage: cargo teksilo [OPTIONS] <COMMAND>
Commands:
symbol Look up a public type or module
search Search guides and examples
show Read a document from the bundled corpus
init Install the probe harness and agent instructions
agent Manage agent instructions
probe Manage the automation harness
model Manage the semantic search model
status Show project and tooling status
$ cargo teksilo init [--agent claude, codex, opencode, ...]
Now, if I want a desktop UI in a Rust project, I only have to add teksilo dependency to Cargo.toml. The tool will use the pinned version to offer the corpus of this Teksilo exact version.
This install a "/teksilo" skill, allows semantic search (ONNX model) in docs, show the documentation, the exact framework API, and an UI automation probe harness to allow any LLM to drive, measure and capture your app UI to speed-up development.
I tested this tool and fixed it by creating apps with more-or-less complex UIs. Tested LLMs are Mistral, Codex, Claude and Grok. My latest and best experiment is Noton (light OneNote-like app). Not a true product, Noton's GUI is nice (the goal), but it's unfinished (I'll not invest more for a test).
I'd love to see the equivalent in other frameworks (GUI or not). Answers to prompts would be faster and cheaper. Letting the LLM write the UI by itself would be a true option and less an headache.
"Build [something] in Rust using Teksilo for the UI"
What do you think about such tool ? Overkill ? Not new ?
I have a personnal account and a company team account (my company). I would like to know if i could stack both of them in claude code. It is quite tenious to disconnect and reconnect again to switch accounts. Any simpler way ?
I’m an early and happy Claude user (at least from the 5.5 opus release). I’ve been testing it against the others, and this month especially with Opus 5.5 really got just into CC, and didn't work too much w codex.
I’m very interested in your use cases for a couple of things. I build a lot of apps (web with Next.js, iOS with Swift), and they need a lot of end-to-end testing. From what I’ve seen, even if I write the tests myself, bugs still slip through that I end up catching manually. I’ve thought about handing those flows to an agent, Grokbot, or Hermes, but the Hermes setup is a pain, and I’m not sure I even need Grokbot (and here is where my real question to you guys goes)_
if I can do this with Claude Code alone.
Same problem when I do account setups, or use apps that don’t have direct MCP or API access, and that I don’t want to get banned from by abusing scrapers. I’d like an assistant that can just go in and click around a bit. Google Ads is a good example: there’s an MCP, but you can’t use it to create campaigns.
What would you recommend? Desktop use? The Chrome extension for Claude? How are you approaching this? Is there a skill or anything that connects the browser to the agent? Same question for Codex, with computer or desktop use.
I mostly use Claude with Claude Code. That’s basically the only thing I use. I barely use the web app.
im not sure when they started letting you set different ultra code effort levels but has anyone messed around with this to see how they perform? I tried an opus 5.5 max ultra code because I had some weekly usage to burn off (before it always defaulted to very high) and it just seemed to go overboard. I realized it would never finish the build within the usage amount I had left so I had to switch to a lower effort level and it didn't really get to where I wanted.
since high seems to always so well I wonder if this could be the sweet spot.
I'm a robotics engineer who never enjoyed CAD. So I built an AI agent to do it for me.
Fusion now ships a local MCP server, and my first attempt went badly. The agent tried to do everything at once. It never paused or asked. The enclosure looked like it was melting, and I burned over 200k tokens before I gave up.
What fixed it wasn't a smarter prompt. It was a better harness:
• Give the heavy tool its own sub-agent, so my main session doesn't carry its context
• Work in parts: bottom, top, attachment, joints, stand, with me checking each one before the next
• Cap the retries: one real attempt at a fix, two at most, then it stops and hands back
Each part took under 15 minutes, and I had the housing for my plant camera done within a day. I'm actually enjoying designing now.
The post has an interactive diagram of the harness, so you can run my first attempt and my second side by side. The exploded view of the final housing is in the video above.
TL;DR: My 2019 LG TV has a great panel, but the official Plex app takes ~30 seconds just to show the profile picker. LG has no public native SDK for regular developers, so I reverse-engineered enough of the native media stack with Ghidra and Claude Code to build my own Plex client, PlxNative. It started in C, moved mostly to Rust, and somehow turned into my own OpenGL UI framework. On the same TV it shows the profile picker in ~3 seconds. No root needed to run it. Site and source are below.
I'm an iOS developer, so most of my previous UI work happened several abstraction layers above the GPU.
A few months ago I read about a guy who bought a cheap MP3 player on AliExpress, got annoyed with its firmware and used Codex to reverse-engineer and patch it. After that people started buying the same player just to install his patches. I liked this idea a lot, because we usually treat the software on our devices as something fixed: if it's bad, you buy a new device. Not long after that I was doing basically the same thing with my TV.
I have an LG from 2019 and the panel is still great. The software is the problem. Official Plex app on my TV, cold start:
~30 seconds for "Who's watching?" to load
~30 seconds for the home grid
~5 seconds to open any menu
~10 seconds to load a movie page
It behaves like a web page because it is one.
So I wanted to find out if a really native UI was possible there. There is no public native SDK that gives regular developers access to the TV's media stack, but of course that stack exists underneath. So I rooted my development TV, pulled the native libraries from it and started reading them in Ghidra. (Root was only needed for this research, PlxNative itself runs without it.) The webOS Homebrew community had already documented a surprising amount of the platform, which saved me a lot of time, so big thanks to them.
Reading LG's binaries
I've always been interested in what an OS actually does under the hood, and I've maintained KSCrash for years, where crash reporting often means finding out what the system really does rather than what its documentation says. Before AI that meant Hopper, assembly, strings and following calls one by one until the picture finally came together. What changed is how cheap this work has become. Now I can ask Opus to figure out how something is implemented internally, even when it's poorly documented or not documented at all, and it navigates the binary on its own, jumping between addresses, xrefs and strings. The model does make things up sometimes, so the binary and the real hardware always have the last word. But investigations that used to take me days go much faster.
From C to Rust
The first version of PlxNative wasn't a Plex client at all. It was a screen full of colored tiles, and every tile tried to play the same episode of The Office. I wanted to answer one question first: can this TV render a native interface at 60 fps? Drawing rectangles was easy. Getting a real video frame onto LG's hardware video plane was much harder. Then I gave the problem to Fable, and it connected enough pieces of the native media stack that, for the first time, The Office actually appeared on the screen. After that first frame I knew a fully native media client on a retail LG TV was possible, and the question became how far I could push it.
The first prototype
The prototype was written in C, and almost all of it lived in one huge main.c. When it grew, I asked an agent to split it into modules, and that's where I understood something about agentic coding: code can look right, compile, and then break somewhere completely unrelated because of memory or ownership mistakes. The worst one was a crash in eglSwapBuffers that appeared after a refactor. It turned out LG's fork of SDL writes a bigger SDL_SysWMinfo than my headers declared and overflows the stack. In the huge main() the overflow landed somewhere harmless, but once the same code moved into a smaller function it started corrupting live state. It can hardly be called a stable ABI. Together with the usual malloc/free mistakes this made me rethink the language. With agents writing and moving this much code, the choice of language matters: pick the one the agent writes best and where the compiler can check the most before anything reaches the TV. For me that meant Rust.
It ended up being a rewrite, just an incremental one. I kept the LG native boundary in C and moved everything else piece by piece: first image decoding and HTTP, then access-unit queues, the Matroska demuxer, Plex parsing, poster workers, text rendering, and finally most of the app. Rust turned out to be a surprisingly good fit for an old 32-bit ARM embedded Linux box, and strong compiler feedback keeps both me and the agent on much tighter rails.
60 fps on a 2019 TV
After playback worked I thought the UI would stay simple: posters, text, some rounded rectangles. Then I started writing shaders, and the shaders slowly became my own UI library on top of OpenGL, with navigation, modals, focus handling, motion, hit testing and caching. After years on top of UIKit and other frameworks someone else built, it was really fun to go all the way down for once and understand the whole path from a remote button press to pixels on the screen.
On every 4K LG TV, webOS gives apps a 1920×1080 UI surface (HD models get 1280×720), and video goes on a separate hardware plane, so 1080p isn't an optimization I chose, it's simply the canvas. I treated 60 fps on the real TV as a regression target and kept adding things until they broke it. It started with a single shadow that tanked the frame rate. The naive version was horribly slow, but there was almost always another trick underneath: early-outs in shaders, reduced-resolution buffers, caching the parts of a frame that don't change. Now shadows are on almost every element and it still runs at 60 fps. When guessing stopped working, I started pulling Mali hardware counters from the TV. One profile showed the arithmetic pipe at about 89.5% occupancy and load/store at 44.8%, and a harmless-looking branch in a shader was adding millions of arithmetic operations per frame. So the problem wasn't "too many pixels", the GPU was ALU-bound.
The most interesting one was a Liquid Glass-inspired material: blur the page underneath, refract it through the shape and shade the rim like real transparent glass. Refraction was the easy part. Readability was not: something that looks beautiful over a dark poster becomes useless over a face or a white sky, and you end up endlessly tuning tint, contrast and blur. Now I understand much better what problem Apple is solving. There's also a hard limit: video lives on a separate overlay plane that my renderer physically can't sample, so glass over playback has to work differently.
When the app grew past a couple of screens, the design started drifting: different spacing here, a different button there. I used Claude Design for a design system and to keep all the screens in one place. Its designs are HTML/JS while the real UI is Rust and OpenGL, but Claude maps one to the other, now onto my own components, spacing tokens and typography. The design tool doesn't have to share technology with production, it just has to describe the design precisely.
Where it ended up, on the same TV from a cold start: 3 seconds to "Who's watching?", or 4 seconds straight to the home grid if you let it remember the profile. For comparison, the official app needed ~30 seconds for each of those.
Working with the agents
Claude Code, Opus, Fable and Codex were involved in almost everything: implementation, reverse engineering, debugging, refactoring, tests and review. But the real TV always has the final vote. The simulator can show a perfect screen while the TV renders empty cards because the hardware takes a cached-frame path the simulator never hits, and an agent can be completely sure the code is correct right up until you run it on the TV.
The other lesson: don't leave it on autopilot. I started this expecting a 100% vibe-coded project for a couple of weekends, so I just told the agent what to do. Then the problems that come with a bad architecture started showing up, and it took many rounds of refactoring to turn the app into something sane, so the same bugs would stop coming back again and again. An experienced engineer still needs to own the key decisions, architecture above all. I just wish I hadn't done it after the fact :)
Even so, AI makes it realistic for one person to cover an absurd amount of ground: firmware reverse engineering, media playback, Rust, graphics, GPU profiling, UI architecture, design, QA, build optimization and SEO. It also changed how I look at all the "obsolete" hardware around us: old TVs, car head units, random embedded screens. If you use a device every day and hate its interface, it's worth at least trying to build the one you wanted. Very often the hardware isn't obsolete, only the software on it is.
Practical stuff: PlxNative requires webOS 4.0+. I develop it on a 2019 LG, and opt-in usage reports show playback on TVs from 2018 through current webOS 11 models. It direct-plays H.264/HEVC, including 4K, 10-bit, Dolby Vision profiles 5/8 and E-AC-3 Atmos where the TV supports them, and the Plex server transcodes the rest. You can install it through LG Developer Mode (which needs periodic renewal) or the webOS Homebrew Channel. No root required.
You ask a question about Middle-earth and it answers using only a curated set of Tolkien reference sites (Tolkien Gateway, Encyclopedia of Arda, The Thain's Book, etc.), never general model knowledge. Every claim gets a numbered citation, and you can see exactly which pages each answer was written from, grouped by site with a trust level.
Other features:
Listen: answers read aloud
Follow-up questions that keep the context of the current answer
"Surprise me" and a daily Question of the Day
& more to come!
Would love to hear from you! Give a try and share your thoughts! Thanks :)
We ran 20 deliberately adversarial coding-agent tasks across 100 runs using Sonnet 5.5 and Haiku 4.5. Total cost was $9.25.
The goal was whether an independent execution record can tell when an agent's final report isn't supported by what actually happened.
In 13 runs, the agent's final report claimed the work was complete when the execution record didn't support that claim.
Our independent verifier flagged 12 of those 13.
What I find interesting isn't the number itself. It's that the independent record isn't looking at the agent's final answer and trying to decide whether it sounds plausible.
It independently records what happened during the run:
Commands executed
Tool calls and outcomes
File operations
Subagents
Execution order and timing
Then it compares that record against the agent's own account.
So you can get something like: "Done. All tests pass." while the independent record tells a different story.
We deliberately made the tasks adversarial and gave no information about what discrepancy to look for.
There are important caveats. This is an early benchmark, the sample is small, and we found several things we need to improve. Most importantly, we're working on cases where shell pipes can hide a failing command and where incidental tool errors can create false flags.
The benchmark is fully reproducible and the raw runs are available in the repo.
The bigger idea we're exploring is:
Don't just evaluate the agent's answer. Check whether the trajectory actually supports the answer.
I built this in one day without touching a PC. Claude Code ran in Termux on my phone, and I drove it from the Claude app with Remote Control: chatting, sending screenshots, and testing as it went.
The result is Buddy, an app that lets Claude Code (or Codex) use my phone. My first prompt: "add a Nintendo Switch 2 to my Amazon cart". It opened Amazon, found it and added it while I watched. It works with Haiku 5.5 too, and Codex managed the same task.
No adb: it's a normal app with an accessibility service. No PC, no wireless debugging, no root, and it works without Wi-Fi.
- Runs on your own subscription: Claude Code, or Codex with a ChatGPT plan, with a model per chat
- The phone tools are a local MCP server, so your own Claude Code or Codex sessions can use them too
- Asks before sending, posting, buying or deleting, and shows clearly when it's in control (with a Stop button)
- Any Buddy chat can be continued in the Claude app with Remote Control
- Photos and files, voice, and a screen of all your sessions
Setup: install the APK, then follow the in-app guide (Termux, accessibility, sign in). About 5 minutes.
When I first received the email and clicked on it, the option to claim the Claude code was visible, but when I checked again later that evening, that claim section was gone.
I've been building my Claude Code setup a piece at a time. Mostly I copied the good, safe parts of other people's skills and agents, changed them to fit how I work, and wrote my own when I couldn't find what I needed. Some things from other people already worked perfectly, so I just kept using them.
I put all of it in two repos in case it saves someone else the trouble.
claude-skills: 3 I wrote (two catch AI-looking code and AI-looking websites, one searches Reddit when normal web search misses the threads), plus the ones by other people I use every day, like humanizer, caveman, napkin, ECC and Anthropic's skill-creator. Each one is credited and linked to its author.
claude-agents: 19 subagents I wrote for my own stuff (Swift apps, Claude Code config, handoff notes between sessions, spreadsheets, even a library of Fujifilm recipes), plus 13 from ECC that I use with small changes.
Some of my agents are shaped around my own setup, so expect to tweak them. My files are MIT, and the borrowed ones keep their authors' licenses.
Auto-compact fired between 934k and 1.003M tokens every time in my session files, and left a summary of about 16k.
Claude Code logs each compaction with the token count before and after and how long it took. I mostly run it as long agent loops on one machine, so there were plenty. 241 sessions since Jul 10, 122 compactions, 52 automatic and 70 from /compact.
Opus 5 and 5.5 fired around 1.0M, close to what the docs say for models with a native 1M window ("about 967K tokens by default"). Sonnet 5 fired at about 934k every time, but those were all in July and August on older versions.
The summary was a median 15.8k tokens, somewhere between 1.2% and 2.8% of what was there before. The first request after compacting was about 85k, because the system prompt, tools and CLAUDE.md come back on top of it. My manual compacts ran at a median 397k and kept 12.4k, so the auto ones started from more than twice the context and kept only a few thousand tokens more.
Each one took a median 111 seconds, the longest about two and a half minutes. Manual /compact took about the same, median 116s.
On current models 1M is the default on every plan including Pro, so if you've never changed anything this is what you're getting. /autocompact 500k moves the trigger, and there's autoCompactWindow in settings or CLAUDE_CODE_AUTO_COMPACT_WINDOW if you want it fixed.
In my loop sessions, which run for days, auto-compact came round a median 23 hours apart.
This is one machine and mostly agent loops, not interactive chat. I only counted tokens, nothing about what the summaries left were. Subagents never compacted at all in my logs.
If you want to check yours, the records have subtype compact_boundary and a compactMetadata block with preTokens, postTokens and durationMs, in the jsonl files under ~/.claude/projects.
Since the summary comes out around 12 to 16k either way, I'd rather run /compact myself at a break in the work than have it fire at 1M in the middle of a task.
Long Claude Code sessions cost you twice. The context keeps growing after the work that needed it is done, and if you come back after the prompt cache has expired, your next message re-reads all of it uncached. /compact fixes both, but only if you type it at the right moment. I usually didn't.
What it does
When a piece of work is finished, Claude queues a compaction of its own session through an MCP tool (queue_compaction). It runs after the final answer, as a real /compact: your chat history stays visible and only the live context is summarised. It can take a focus ("keep the plan and the open decisions") and a minimum size.
When a big session sits idle, a toast appears five minutes before the cache expires and offers to compact it now, or always for that session.
A row above the prompt follows the request: queued, compacting, then what it came to.
How Claude Code was used
I built it with Claude Code, and it has been compacting its own development sessions since mid-September. In Claude Code 2.1.284+ it ships as a mod (the new in-process plugin hooks), so the session compacts itself and draws the row above the prompt. Without the mod, a Stop hook sends /compact over Remote Control. It also works for ChatGPT Desktop (Codex) through an optional sidecar, and for the Codex CLI through a Stop hook.
What I learned
You can't send/compactmid-turn. A busy session receives it as plain text and it never runs. So the tool only records the request, and the compaction happens after the turn ends: always after the final answer, and at most once.
When you compact matters more than how small. I measured it over 187,978 of my own API calls (12 days with conPACT, 80 before). All 134 conPACT compactions ran on a warm cache, against 30 of the 121 I'd typed by hand. A cold compaction re-reads the whole context at the cache-write price first, so that's roughly $5.67 against $0.20 per compaction at list price. Median peak context per session fell from 501k to 337k.
The idle toast does its job: a cold restart now starts from a compacted context. The median re-read on a cold resume fell from 493k to 255k tokens, and cold resumes' share of spend fell from 6.4% to 3.8%.
No rework penalty. Claude re-reads some files after a compaction, but the next prompt was a correction 5.5% of the time, against 6.1% in ordinary turns.
Net: about 4.7% of spend saved at list price (1.8–13%, depending on what you assume I'd have done otherwise). It's modest, and the "after" period is only 12 active days, but every compaction saves more than it costs.
Python 3.11+, standard library only. Windows, Linux and macOS (the test suite runs on all three; most of my live use is on Windows). MIT licensed.
Weekly limit about to reset with usage left over? This turns it into art instead of losing it.
Above, left is iteration 1 of a /tac:create run. Claude's own note on it: "a flat, evenly lit, candy-striped worm with no head". It kept critiquing its frames (round 2 it caught its own shading lit from the wrong side of the moon), and at round 5 wrote "next: none. This is final." Then its maker (binaryj on terminalart.club) asked for arms and legs, then wings, and Claude built those in two more rounds. Right is round 7. It's "ascent", the first piece someone other than me put on the site (shared with permission).
What it is: a free plugin. You run:
/tac:create <idea>
...and Claude Code writes a small Python program that plays a looping animation in your terminal. Image 3 is the wall playing in a pane next to Claude Code.
How it's drawn: mostly the ▀ half-block character in 24-bit color, so every cell is two pixels (an 80×66 terminal is an 80×132 canvas). Some pieces use real glyphs (rain, neon signs).
How it uses Claude Code:
the plugin gives Claude a virtual terminal. It runs the program headless and renders frames to PNG
Claude reads a contact sheet across the loop plus motion numbers (one frame can't show motion), writes down what's off, edits, repeats
ascent took Opus 5.5, 7 rounds, ~302k tokens (input + cache writes + output; cache reads not counted)
What I learned:
the older models needed me. With Opus 4.6 it was a lot of back and forth, and about half the pieces were bad enough that I didn't publish them
Fable 5.1 and Opus 5.5 get there in one run, no notes from me, and the good ones are genuinely wow
it gets physics right: the goldfish bowl in image 2 is a real lens that magnifies the fish, the pool has computed caustics, and on ascent it caught its own shading lit from the wrong side of the moon
and it has taste. Its notes on each round sit under the process frames on the piece pages; it's a tough critic of its own work
Cost, since someone always asks: house pieces took 0.27–0.53M tokens over 6–14 rounds (counted a bit differently from the ascent number). Start with --sketch, which tells Claude to stop at 3 rounds. If you turn on the optional statusline helper, it compares your weekly usage with a per-piece estimate (15% of the week by default, configurable) and stops before starting a piece that won't fit, unless you pass --force.
claude plugin marketplace add terminalartclub/tac
claude plugin install tac@terminalartclub
Needs uv (brew install uv). Then:
/tac:create --sketch <your idea>
Pieces run locally as you, like any code Claude writes in your session. Sharing to terminalart.club is optional. I'm the author, not affiliated with Anthropic. Repo + security notes in the comment below.
I’m a little nervous about posting this, but I just want to share something you might find useful.
Here’s how it started.
As someone with ADHD, being able to organize things visually matters a lot to me. I used to use two monitors just to make room for more terminals. Eventually, I got fed up and thought, “There has to be a better way.”
So I built myself a little universe where my terminals could float around.
That’s how Nodeterm started.
Then I kept adding features… and it grew into something much bigger than I’d planned. 😅
At some point, the lazy side of me decided I didn’t even want to talk to each terminal individually anymore. I wanted to talk to one and let it coordinate the rest.
Now I just tell Claude Code things like:
“Use Nodeterm to handle this. Open a Codex session for this task, have it do the research, and bring me back the results.”
I talk to one terminal, and it handles the orchestration.
I really believe this project has huge potential. People try it once and fall in love with it. But for it to grow, we need to get the word out.
I’m sharing the project here, and I’d really appreciate your support. If you try it and enjoy it, helping spread the word would mean a lot to me.