Anthropic released Claude Fable 5.1 on September 1, positioning it as the company’s most capable and cost-efficient model for complex reasoning tasks. The headline number: 90% coverage on the ARC-AGI-2 benchmark at 32% less cost per task than its predecessor, Fable 5, which carried a verified price tag of $5.45 per task.

The benchmark gains in context

Fable 5.1’s 90% coverage on ARC-AGI-2 represents a notable step up from Fable 5’s verified score of 89.2%. On Terminal-Bench 4.0, Fable 5.1 scored 55.8%, compared to Fable 5’s 42.0%. That’s a 13.8 percentage point jump, roughly a 33% improvement in raw scoring terms.

Terminal-Bench-Science 0.1, a recently introduced evaluation for scientific reasoning, told an even more striking story. Fable 5.1 posted 52.6% versus Fable 5’s 24.7%. The model more than doubled its predecessor’s performance on scientific tasks.

Fable 5.1 also topped multiple intelligence and agentic indexes for 2026, including an Artificial Analysis score of 66 under maximum effort conditions. The model ships with a 1 million token context window, 128K maximum output, and an adaptive reasoning system featuring five distinct effort levels.