Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning.

Anthropic's Claude leads enterprise coding benchmarks while OpenAI's GPT-5.6 wins on speed. What the AI arms race means for tech and crypto

Anthropic's new everyday model Claude Opus 5 undercuts its own frontier product on cost and beats it almost everywhere that counts.