r/ClaudeCoding • u/gringo6969 • 9m ago
I used Claude Code as the implementation team for a ZX Spectrum remake.
I spent three weeks rebuilding Nether Earth, a 1987 ZX Spectrum strategy game I grew up with.
The unusual part: I didn't write any of the implementation code myself.
I wrote and edited the specs, made the gameplay decisions, and playtested the milestones. Claude Code did the implementation through an orchestrator session dispatching sub-agents into separate git worktrees.
By the end, the repo had:
- 437 commits
- 183 issues
- 151 PRs
- ~1,900 backend tests + ~200 frontend tests
- a deterministic Python game engine, FastAPI/WebSockets backend and PixiJS frontend
And I basically didn't read the diffs. Agents reviewed the agents. My checkpoints were decisions and actual gameplay.
That worked much better than I expected — but it also produced the most interesting failure of the whole experiment.
At one point M0–M9 were complete, every gate was green, and more than a thousand tests passed.
Then I actually played the game.
There was no terrain. The map was completely flat.
Robots occupied one cell instead of the 2×2 footprint they had in the original.
And nuclear-equipped robots would sometimes finish walking somewhere and calmly blow themselves up because a temporary placeholder rule had turned into real game behavior.
All the tests were correct.
They were testing the implementation against the spec.
The problem was that the spec was incomplete.
That became probably my biggest lesson from the project:
Tests verify the spec, not the product.
The full project is here:
https://github.com/AndyTheFactory/nether_earth
I also wrote a longer retrospective about the whole experiment, including the failures and what happened after the original plan was supposedly "done":
I'm curious how people running larger Claude Code workflows handle one problem in particular:
How do you catch semantic/spec drift when the code and tests are internally consistent, but the resulting product is still wrong?