OpenAI published an audit of SWE-Bench Pro on July 8, 2026 and estimated that roughly 30% of its tasks are broken. The reported issues make a familiar leaderboard assumption unsafe: every task in the denominator is a valid, equally interpretable trial.
Primary source: OpenAI, “Separating signal from noise in coding evaluations”.
The operational response should not be “ignore all benchmarks.” It should be: version task validity, preserve disputed cases, and publish how conclusions change across plausible denominators.
Model task state separately from model result
task validity: unreviewed | valid | broken | disputed






