A green test suite is not real evidence. It is often a closed argument loop. The same agent wrote both code and checks.
Freeze an oracle before any agent run. Then let every patch fail in public. Cheap tokens do not weaken this rule.
Take a side
Stop treating generated tests as quality control. A model that authors both sides grades itself. That process is narrative, not verification.
Retry-heavy coding loops make the narrative cheaper. They also make the story smoother. Smooth output is the actual danger here.






