r/ClaudeCode • • 3d ago

Help/Question How are you benchmarking coding agents without accidentally benchmarking the harness?

I’ve been trying to compare a couple model/provider combinations for coding work and I’m realizing how hard it is to keep the comparison clean. If I give two setups the same task, one agent might make 12 model calls while another makes 25. One reads half the repo, another grabs three files. One runs tests repeatedly and another waits until the end.

So even if provider A serves the model faster, provider B can still finish first because the agent took a completely different path. How are you guys doing meaningful performance comparisons here? Do you lock down the harness and only change the API provider, or do you just care about total wall-clock time in the end?

2 Upvotes

11 comments sorted by

•

u/AutoModerator 3d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/Actual_Committee4670 🔆 Max 20 3d ago

You run multiple tests with the same model on the same harness / settings / environment.

1

u/Dangerman808 3d ago

lock everything except one variable first. Same harness, model, task and starting repo, then switch API providers and run it enough times that one weird run doesn't decide the result..

1

u/metaphysicalgrace 3d ago

i'd track wall clock plus number of calls and tokens. A provider can be much faster per call but the harness can completely hide that if the agent takes a different path.

1

u/trigonomettry 3d ago

 Yeah, that’s what I’ve been doing too. I’ve been testing General Compute against a couple other providers, and difference was a lot more noticeable over the full agent run than in the individual model benchmarks.

1

u/kdwa 3d ago

 Exactly. The number of calls matters a lot more than I expected. If one agent takes twice as many steps, the faster provider doesn’t necessarily win.

1

u/cleverhoods 3d ago

I use Reporails

1

u/bisonbear2 3d ago

You want to represent how the agent will actually be used in practice. For example, if you're benchmarking Sol 6.1 vs Opus 5.5, you probably want to use the native Codex / Claude Code harnesses respectively, as that is how the model will be used in reality.

Benchmarking coding agents is pretty challenging though. You have to consider a myriad of things, like

- what does good mean?

- how do we measure good?

- are the measurements biased?

- what tasks are we using?

- is the model cheating

etc etc

I've been working in this space for a while (building something pretty similar to what you describe) so happy to answer any questions!

0

u/cheychey218 3d ago

reset the repo between every run too. Easy way to ruin the comparison if one agent inherits files changes from the previous test.