OpenAI's audit reveals 30% of tasks in SWE-Bench Pro are flawed, prompting the company to withdraw its recommendation for replacing SWE-bench Verified with this benchmark.

OpenAI's audit reveals 30% of tasks in SWE-Bench Pro are flawed, prompting the company to withdraw its recommendation for replacing SWE-bench Verified with this benchmark.

OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its…