OpenAI published an audit of SWE-Bench Pro on July 8, 2026 and estimated that roughly 30% of its tasks are broken. The reported issues make a familiar leaderboard assumption unsafe: every task in the denominator is a valid, equally interpretable trial.

Primary source: OpenAI, “Separating signal from noise in coding evaluations”.

The operational response should not be “ignore all benchmarks.” It should be: version task validity, preserve disputed cases, and publish how conclusions change across plausible denominators.

Model task state separately from model result

task validity: unreviewed | valid | broken | disputed