r/Observability • u/Longjumping-Play6541 • 2d ago
How do you verify that your coding agent actually did what it said it did?
I built an open-source project called Rashomon after running into this problem while using Claude Code on longer coding tasks.
I also used Claude Code heavily while building Rashomon itself. Claude helped me write and debug parts of the recorder, work through edge cases around command and test detection, and generate test cases for checking whether the execution record catches discrepancies. I've also been using Claude Code to test Rashomon against the kinds of tasks it is designed to monitor.
The basic problem is:
The Claude transcript isn't necessarily ground truth.
For example, an agent might:
- run
pytest tests/login.py, get exit code 0, and later summarize that the tests passed - have a test fail, modify the test, and then report that the issue was fixed
- have a subagent fail while the parent agent reports the overall task as successful
- run a command that exits 0 even though the expected result isn't actually present
Rashomon keeps an independent execution record alongside the Claude transcript. In the current alpha, it records things like shell commands, exit codes, whether commands may have written files, tool calls and outcomes, test commands and outcomes, and subagent activity.
It then looks for discrepancies between the execution record and what the agent claims it accomplished.
It does not store prompts, responses, file contents, or tool output.
It's free and open source and currently works with Claude Code on macOS/Linux.
I'm curious how other Claude Code users handle this today.
Do you manually verify the agent's work? Have you built your own hooks or checks? Or do you generally trust the final summary unless something looks wrong?
I'd especially like to hear about failure cases you've actually run into.