A coding-agent pass rate without a published retry cap is an incomplete measurement rather than a fair comparison. Hidden extra attempts inflate success in the same quiet way extra minutes inflate a closed-book exam score. Honest reporting therefore treats retry policy, wall-clock limits, and tool-trace identity as first-class dataset fields. Scores that omit those fields should be read as anecdotes rather than as comparative evidence.

Developers currently compare agentic coding tools with a single percentage that looks decisive in a README table. That number usually collapses several invisible knobs, including how many times the agent may regenerate a patch after a red test. Two systems can share a pass rate while one spent a single attempt and the other burned a silent loop of repairs. A methodology that does not freeze those knobs is closer to marketing copy than to a controlled experiment.

The proposed dataset is a sealed task pack rather than a live repository that changes under the runner. Each task record should include a unique identifier, a frozen prompt template, a hidden test digest, and an uneditable oracle command. The pack also needs an explicit retry ceiling and a wall-clock ceiling so that time and persistence are not left as folklore. Without those ceilings, later readers cannot reconstruct the labor that produced the headline success rate.