A claim circulating on social media that OpenAI’s GPT-6 Astra scored 98.6% on the ARC-AGI-3 benchmark would, if true, represent one of the most significant leaps in AI capability ever recorded. The problem: there’s no verified evidence to back it up.

The ARC-AGI-3 benchmark, which launched on March 25, 2026, is specifically designed to test whether AI models can navigate interactive environments without instructions or predefined objectives. When models first encountered the benchmark, scores came in below 1%.

What the leaderboard actually shows

The current top performer on ARC-AGI-3 is Anthropic’s Claude Opus 5, sitting at roughly 30.2%.

OpenAI’s own GPT-5.6 Sol, its most recent model with verified benchmark results, managed 7.78% under official testing conditions. When OpenAI used its custom Responses API settings, that number climbed to 38.3%.