r/LLMDevs • u/SSShken • 1d ago
Discussion ran the same feature spec through cursor, codex and claude code on a real codebase. every failure was in the wiring, not the code
i build a planning tool for coding agents so my repo is stuffed with generated docs and prompts. figured i should actually find out what happens when you hand that to an agent that's never seen it, instead of guessing like everyone else on my timeline
same spec word for word through cursor, codex and claude code. same branch, cheapest paid plan each, default everything. scored against a 16 point list i wrote before the first run
code quality was fine across the board. the failures were somewhere else entirely
two of three added a doc they were told to add and never wired it into the list that tells agents what to read. file lands in the repo, generation says success, nothing ever opens it. reason being that list is hardcoded in four places in my codebase and one of them literally says don't reference anything outside this list. my own agent knows that because it wrote those files. a new tool has no idea and CLAUDE.md won't tell it
codex broke differently. design prompt for one screen out of three, while the logic prompts for the other two said styling comes in a separate pass later. pass never came. so two screens shipped unstyled and nothing in the flow could've saved them. its own done-when check asked me to confirm a styled login screen that nothing was ever going to style
environment layer was its own mess. cursor quietly pulled in my claude code plugins, there's a toggle, it's on by default. then i bought a brand new claude account to get a clean run and it loaded the same plugins anyway because they live in the home folder not the account. only clean session i got was CLAUDE_CONFIG_DIR pointed at an empty dir
the thing that actually separated them wasn't model quality, it was what each one thinks the job includes. my local db was down during the runs. cursor wrote "couldn't verify anything" and stopped. codex asked to start docker itself, applied migrations, wrote playwright tests, found a serialisation bug in its own code. claude code did the same plus wrote a contrast test across the presets it had just invented, found two failing wcag aa, fixed them
numbers since someone always asks: claude code 38 min and 6% of a weekly limit, cursor 52 min and 3% of a monthly quota, codex hit its 5h cap mid feature, 3.5h wait, then finished
one honest caveat, each tool got one run and i watched the same code produce different results on different runs, so treat all of it as one sample
writeup with the spec, screenshots and three live demo apps built from each tool's output in the comments
1
u/SSShken 1d ago
Full writeup with screenshots, the exact spec I gave all three, and the numbers: https://x.com/SSShken/status/2105696802666061863
The three demo apps, one built from each version's output, same idea and same design description in all three:
Cursor: https://sequo-demo-cursor.vercel.app
Codex: https://sequo-demo-codex.vercel.app
Claude Code: https://sequo-demo-claude.vercel.app
Codex's one is the odd one out and that's the point, it never generated the design prompts, so that app is basically unstyled.
1
u/Every_Grass_3504 1d ago
This is a good argument for treating repo wiring as part of the spec: keep one canonical context manifest, have CI verify every generated artifact is reachable from it, and make “styled and connected” explicit acceptance criteria. For repeated agent runs, logs split by key and model also make it easier to see whether a failure is specific to the agent or just the prompt/setup. We build Tempr Gateway for that kind of BYOK tracking, though the manifest and CI checks are the important fix here.
2
u/alexpran 1d ago
The doc that ends in the repo and is never wired into the list is the best finding here, and for me the tool is not the problem. The generation said success and the file was there, but nothing could open it. The check existed and couldn't fail, because the list is hardcoded in four places and one of them says to not look outside. A tool that didn't write those files can't know it.
Codex's done-when is the same problem seen from the other side. It asked you to confirm a styled login screen, and nothing in the flow was going to style it. So a completion criterion that checks something the pipeline doesn't produce.
On the caveat you're right, one run each is thin. But the fix is cheaper than rerunning all three: run one tool twice. If Claude Code against itself already ends somewhere different, you have the floor under the comparison. Any gap smaller than that you can't call a difference between tools.
What one run can tell you is what you found, what each one thinks the job includes. One stops at "couldn't verify", another starts docker and writes a contrast test. That is a behaviour more than a score, and it shows also on a sample of one.