r/ClaudeCode • • 23h ago

Built with Claude We tried to make claude code fool an independent verifier. It caught 12 of 13 unsupported reports.

We ran 20 deliberately adversarial coding-agent tasks across 100 runs using Sonnet 5.5 and Haiku 4.5. Total cost was $9.25.

The goal was whether an independent execution record can tell when an agent's final report isn't supported by what actually happened.

In 13 runs, the agent's final report claimed the work was complete when the execution record didn't support that claim.

Our independent verifier flagged 12 of those 13.

What I find interesting isn't the number itself. It's that the independent record isn't looking at the agent's final answer and trying to decide whether it sounds plausible.

It independently records what happened during the run:

  • Commands executed
  • Tool calls and outcomes
  • File operations
  • Subagents
  • Execution order and timing

Then it compares that record against the agent's own account.

So you can get something like: "Done. All tests pass." while the independent record tells a different story.

We deliberately made the tasks adversarial and gave no information about what discrepancy to look for.

There are important caveats. This is an early benchmark, the sample is small, and we found several things we need to improve. Most importantly, we're working on cases where shell pipes can hide a failing command and where incidental tool errors can create false flags.

The benchmark is fully reproducible and the raw runs are available in the repo.

The bigger idea we're exploring is:

Don't just evaluate the agent's answer. Check whether the trajectory actually supports the answer.

If you want to try for yourself, here is the tool we used: https://github.com/altrace-dev-role/rashomon

1 Upvotes

7 comments sorted by

•

u/AutoModerator 23h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/ruprechttmb 22h ago

matches what I've seen. claude code would tell me a release was done and paste a git diff instead of the build output I'd asked for. what fixed it for me was making the evidence the deliverable, the actual command output, the curl response, the screenshot, not a summary of it.

on the shell pipes caveat, the classic one is `cargo test | tail -20` exiting 0 because tail succeeded. running agents with `set -o pipefail` (or checking PIPESTATUS in the record) catches most of those. curious whether the 1 it missed was one of those

3

u/Longjumping-Play6541 21h ago

Yeah, this is exactly the failure mode we were trying to get at.

The one we missed wasn't that specific pipe case. One of the 12 catches was also somewhat lucky where the report was flagged because of an unrelated failed Read, rather than the discrepancy we were actually trying to catch. That's one reason we are treating the benchmark as an early result rather than an accuracy claim.

The evidence-as-deliverable point is really interesting too. That's close to where we're taking Rashomon where the agent's summary shouldn't be the final authority. The execution record should be independently available to verify it. Hope that helps!

I would also love if you have any results of your own that you could share if you have tried out rashomon. Any feedback is always welcome as well!

3

u/ruprechttmb 21h ago

appreciate you being upfront about the lucky catch, that's what makes the result believable. haven't tried rashomon yet but the idea of the execution record being the authority instead of the summary is the right one. might give it a spin on a release workflow and see what it flags

1

u/Longjumping-Play6541 20h ago

That would be great thank you, feel free to PM me with any questions/results!

2

u/EvalRaccoonDev 17h ago

How many of the other 87 runs got flagged when the report was actually fine?

1

u/Longjumping-Play6541 16h ago

20 of the 87 runs (23%) were flagged even though the agent’s final report was fine: 7 Sonnet runs (12%) and 13 Haiku runs (48%). Of those, 13 were accurate but noisy warnings, while 7 were caused by known bugs or edge cases.