Nearly 30% of SWE-Bench Pro tasks are flawed
Jul 09, 2026
01:04 pm
What's the storyOpenAI has raised concerns over SWE-Bench Pro, a popular coding benchmark.
The company conducted an audit of the test and found that nearly 30% of its tasks are flawed in ways that misrepresent what an AI model can actually do.









