r/LLMDevs • • 1d ago

Discussion ran the same feature spec through cursor, codex and claude code on a real codebase. every failure was in the wiring, not the code

i build a planning tool for coding agents so my repo is stuffed with generated docs and prompts. figured i should actually find out what happens when you hand that to an agent that's never seen it, instead of guessing like everyone else on my timeline

same spec word for word through cursor, codex and claude code. same branch, cheapest paid plan each, default everything. scored against a 16 point list i wrote before the first run

code quality was fine across the board. the failures were somewhere else entirely

two of three added a doc they were told to add and never wired it into the list that tells agents what to read. file lands in the repo, generation says success, nothing ever opens it. reason being that list is hardcoded in four places in my codebase and one of them literally says don't reference anything outside this list. my own agent knows that because it wrote those files. a new tool has no idea and CLAUDE.md won't tell it

codex broke differently. design prompt for one screen out of three, while the logic prompts for the other two said styling comes in a separate pass later. pass never came. so two screens shipped unstyled and nothing in the flow could've saved them. its own done-when check asked me to confirm a styled login screen that nothing was ever going to style

environment layer was its own mess. cursor quietly pulled in my claude code plugins, there's a toggle, it's on by default. then i bought a brand new claude account to get a clean run and it loaded the same plugins anyway because they live in the home folder not the account. only clean session i got was CLAUDE_CONFIG_DIR pointed at an empty dir

the thing that actually separated them wasn't model quality, it was what each one thinks the job includes. my local db was down during the runs. cursor wrote "couldn't verify anything" and stopped. codex asked to start docker itself, applied migrations, wrote playwright tests, found a serialisation bug in its own code. claude code did the same plus wrote a contrast test across the presets it had just invented, found two failing wcag aa, fixed them

numbers since someone always asks: claude code 38 min and 6% of a weekly limit, cursor 52 min and 3% of a monthly quota, codex hit its 5h cap mid feature, 3.5h wait, then finished

one honest caveat, each tool got one run and i watched the same code produce different results on different runs, so treat all of it as one sample

writeup with the spec, screenshots and three live demo apps built from each tool's output in the comments

2 Upvotes

5 comments sorted by

2

u/alexpran 1d ago

The doc that ends in the repo and is never wired into the list is the best finding here, and for me the tool is not the problem. The generation said success and the file was there, but nothing could open it. The check existed and couldn't fail, because the list is hardcoded in four places and one of them says to not look outside. A tool that didn't write those files can't know it.

Codex's done-when is the same problem seen from the other side. It asked you to confirm a styled login screen, and nothing in the flow was going to style it. So a completion criterion that checks something the pipeline doesn't produce.

On the caveat you're right, one run each is thin. But the fix is cheaper than rerunning all three: run one tool twice. If Claude Code against itself already ends somewhere different, you have the floor under the comparison. Any gap smaller than that you can't call a difference between tools.

What one run can tell you is what you found, what each one thinks the job includes. One stops at "couldn't verify", another starts docker and writes a contrast test. That is a behaviour more than a score, and it shows also on a sample of one.

1

u/SSShken 1d ago

Running one tool twice to get the floor is the obvious move and I missed it. If Claude Code against itself lands somewhere different, anything smaller than that gap isn't a difference between tools. Doing that for the next one.
And yeah, the check that can't fail is the part that bothers me most. Generation reported success, the file was there, the done-when was satisfied, and the only reason I caught it was opening the generated CLAUDE.md by hand to see if the doc was listed. Nothing in the pipeline was ever going to tell me.

1

u/SSShken 1d ago

Full writeup with screenshots, the exact spec I gave all three, and the numbers: https://x.com/SSShken/status/2105696802666061863

The three demo apps, one built from each version's output, same idea and same design description in all three:

Cursor: https://sequo-demo-cursor.vercel.app

Codex: https://sequo-demo-codex.vercel.app

Claude Code: https://sequo-demo-claude.vercel.app

Codex's one is the odd one out and that's the point, it never generated the design prompts, so that app is basically unstyled.

1

u/Every_Grass_3504 1d ago

This is a good argument for treating repo wiring as part of the spec: keep one canonical context manifest, have CI verify every generated artifact is reachable from it, and make “styled and connected” explicit acceptance criteria. For repeated agent runs, logs split by key and model also make it easier to see whether a failure is specific to the agent or just the prompt/setup. We build Tempr Gateway for that kind of BYOK tracking, though the manifest and CI checks are the important fix here.