A coding-agent percentage remains marketing until protocol, canaries, and a cost ledger freeze beside it. Teams still quote a lonely pass rate as if that number could travel without a suitcase of hidden choices. Prompt text, tool allowlists, sandbox images, and grader prompts often move the score more than the model does. The honest unit of publication is therefore a registered protocol rather than a percentage standing alone on a slide.

The protocol resembles a flight plan much more than a souvenir photograph of cruising altitude. A photograph of thirty thousand feet says nothing about fuel load, weather, or whether the altimeter was calibrated. Agent leaderboards behave the same way when they publish altitude without publishing the plan that produced it. Pre-registration does not make models stronger; it only makes later numbers comparable to the original claim.

A useful registry records the corpus identifier, the hidden holdout rule, and the grader isolation boundary before scored runs. It also records decoding settings, the tool allowlist, and the network policy that the sandbox will enforce during execution. Those fields look bureaucratic until a rerun drifts after someone quietly upgrades a linter inside the evaluation image. The registry exists to make that drift visible to readers, not to decorate a repository README with extra YAML.