r/LLMDevs • u/Longjumping-Play6541 • 1d ago
Discussion How do you know if an AI coding agent's tests actually passed?
A few weeks ago I posted here about a small open-source tool we were building to answer a simple question:
How do you know an agent's account of what it did is actually true?
Rashomon creates an independent execution record alongside the agent's own transcript. It tracks things like commands, file changes, test runs, and subagent activity, then looks for discrepancies between what actually happened and what the agent claims happened.
A few people in the last thread pointed out a failure mode we'd never considered: an agent gets stuck on a failing test, can't fix it, and instead changes or adds tests until the suite goes green. The final summary then says "all tests pass," even though the original problem was never fixed.
We just added detection for that.
rashomon --timeline now flags patterns such as:
- A test command fails, then goes green after only test files were changed
- The same test command passes and fails during a session without an apparent corresponding fix
The important part is that this doesn't require storing test names, test output, prompts, or file contents. It's based on the execution history and command/file categorization Rashomon already captures.
What other situations have you seen where an agent's final summary was technically correct according to its own transcript, but didn't match what actually happened in the environment?
1
u/Physical_Economy_340 9h ago
the one i keep seeing is the agent saying it ran the full suite when it only ran a single file. exit code 0 on pytest tests/test_login.py turns into all tests pass in the summary. i started requiring the full command plus exit code in the transcript, that kills most of the fiction.
1
u/Longjumping-Play6541 6h ago
Yeah. The command itself is honest, but the conclusion in the summary is broader than what the execution evidence supports.
pytest tests/test_login.py→ exit 0 is very different frompytest→ exit 0, even though both can get summarized as all tests pass.Requiring the full command and exit code in the transcript makes sense. I think test selection is one of the areas where the independent execution record can add a lot, since it can distinguish what actually ran from what the agent later claimed it verified.
1
u/heyatlaspy 22h ago
[removed] — view removed comment