Two API settings. Same model. Same benchmark. Same task set.

13.3% to 38.3%, using one sixth the output tokens.

OpenAI published that result about GPT-5.6 Sol on ARC-AGI-3, and it is the cleanest natural experiment the field has produced on a question I have been arguing from first principles for a year. Nothing about the model changed. Everything that changed was around it.

The score is Relative Human Action Efficiency, not a pass rate, and OpenAI estimates the average human tester at 48% on the same set. So the model went from 34.7 RHAE points behind a human to 9.7 behind, and the entire move came from configuration. By my arithmetic that is about 72% of a gap people had been attributing to the model.

There is a blunter version of the same fact. On the public leaderboard for one of these games, no frontier model gets past the first level. With the reconfigured harness, GPT-5.6 Sol solves all six.