A coding-agent percentage is a measurement only when a protocol file exists before the first task runs. Screenshots, live demos, and leaderboard rows can look precise while the dataset and the scorer still move. Reviewers should treat an unpublished protocol like a lab treats an unlabeled reagent on a crowded bench. The reading may be physically real, and it still cannot be compared with any later reading.
The working analogy is a kitchen scale that is zeroed after the flour is already in the bowl. Moving the tare after the pour does not invent extra grain, yet it quietly rewrites every later claim. Agent evaluations repeat that gesture when tasks are added because a model looked weak on Tuesday. A frozen protocol does not make the agent smarter; it makes tomorrow's percentage talk about the same object.
A useful dataset for this method is a holdout corpus with an explicit split, not a folder of favorite bugs. Each task needs a stable identifier, a language, a fixture path, a timeout, and a public-versus-hidden test flag. Tasks that authors used while prompting the agent belong in a development split and must never enter the reported score. The reported split should stay sealed until the protocol document is written, because curiosity itself is a contamination path.






