Your coding agent finishes a task and tells you the tests passed.
You accept the report. Later, when one claim does not match what happened, you go back through the log to find where the story diverged - and the log cannot tell you. It records what the agent said. It does not record what ran.
That gap is not solved by making the agent's report more persuasive. A report is still a report unless the run is backed by execution evidence.
The question underneath
"Did the tests pass" is a question about an outcome. The useful question is one level down:






