Update:

ARC Prize co-founder François Chollet responded to OpenAI's results by distinguishing between two kinds of test setups. Harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits, he said. General-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game. In effect, Chollet is conceding that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage.

He noted that ARC Prize has had "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," and welcomed the company "starting to figure out the answer." Different providers using different settings does create "a potential parity issue," Chollet said, but he considers that acceptable "as long as the settings and the cost are clearly reported."

Original article:

OpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmark, OpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent.