Anthropic’s Claude Sonnet 5 arrived on June 30, 2026, and immediately made a case for itself where it counts most: not on a benchmark cooked up in a lab, but on real tasks, run by real users, graded on whether the thing actually worked.
The model landed at #6 overall on the Agent Arena leaderboard, posting a net improvement score of 7.38% plus or minus 1.30%. That number reflects how much better users fared with Sonnet 5 compared to competing models across more than 1 million live sessions tracked by the platform.
What the numbers actually mean
Claude Fable 5 (High) holds the top spot, with several models from OpenAI’s GPT-5.5 series and Opus variants occupying the middle ground.
Where Sonnet 5 distinguishes itself is in the confirmed success rate category, where it ranks second overall at 12.24%. In plain terms: when users set it a task, it completed that task at a rate that outpaced nearly every other model on the board.






