OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark.

OpenAI's audit reveals 30% of tasks in SWE-Bench Pro are flawed, prompting the company to withdraw its recommendation for replacing SWE-bench Verified with this benchmark.

OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its…