A coding-agent score that hides retries, timeouts, and helper calls is closer to marketing than to measurement. Pass rate remains useful, but it becomes comparable only after an attempt budget and a retry ledger sit beside it. This methodology treats those controls as first-class fields, so a free-tier lab can be reproduced without inflating a lucky run. The numbers that survive this process are smaller, slower to collect, and far easier to defend.

Public talk about agentic coding still gravitates toward a single green check on a pull request. That check rarely records how many regenerations, tool loops, or stretched timeouts stood behind it. Two labs can therefore publish identical pass rates while spending very different amounts of model effort. A methodology that cannot show the spend is not a benchmark; it is a press release with a denominator.

Freeze the dataset before the first call

The proposed dataset is a small stratified bundle, built on purpose rather than scraped from a contest dump. Each task carries a frozen fixture directory, a public oracle, a wall-clock timeout, and a hard cap of k independent attempts. Stratification should mix repair, greenfield, and regression work so one easy cluster cannot dominate the mean. Tasks that leak the oracle into the prompt belong in a quarantine file, not in the scored set.