Nearly 30% of SWE-Bench Pro tasks are flawed

Jul 09, 2026

01:04 pm

What's the storyOpenAI has raised concerns over SWE-Bench Pro, a popular coding benchmark.

The company conducted an audit of the test and found that nearly 30% of its tasks are flawed in ways that misrepresent what an AI model can actually do.