OpenAI changed two API settings on ARC-AGI-3 and GPT-5.6 Sol went from 13.3% to 38.3% while spending six times fewer output tokens. Same model, same weights, same benchmark.

ARC Prize scored GPT-6 Astra at 62.7% on ARC-AGI-3 with its own harness and 99.9% with OpenAI's, then revised five metrics after launch.

OpenAI changed two API settings on ARC-AGI-3 and GPT-5.6 Sol went from 13.3% to 38.3% while spending six times fewer output tokens. Same model, same weights, same benchmark.