Grok 4.5 scores 91.3% on VulcanBench, outperforming Claude Fable 5 and GPT-5.6 Sol on real-world coding tasks with lower per-task costs.

Grok 4.5 leads the SWE Marathon benchmark at 29.0%, beating Claude Opus 4.8 and Fable while offering $2 per million token API pricing.

Grok 4.5 scores 91.3% on VulcanBench, outperforming Claude Fable 5 and GPT-5.6 Sol on real-world coding tasks with lower per-task costs.