Originally published on hexisteme notes.
My agent handed me a recommendation that was the exact opposite of a warning sitting in a file it had created four days earlier, specifically to stop this from happening again.
The setup: I keep an AI coding agent working on a competition entry, under a written contract document it is supposed to follow. Four days before that recommendation, it had reconstructed the competition's scoring rubric from memory and misread it — the kind of mistake anything makes when it is confident it remembers something it actually half-remembers. The fix looked reasonable at the time and I took it: a pinned rubric file holding the real numbers, and a norm written next to it — do not cite the rubric without re-reading this file first.
Four days later it made the same call again. Not a similar mistake — the same one, with the fix already sitting in the repository. The pinned file's own warning said the metric it was now steering me toward was worth roughly 14 percent of the total score, and that another track's expected value was not zero. It had written that sentence itself. Then it handed me the opposite as a load-bearing recommendation.
Worse than a soft constraint






