Claude Opus 5 and Claude Opus 4.8 bill the same $5 per million input tokens and $25 per million output, and on identical prompts the default Opus 5 configuration cost 3.1x more. The reason is adaptive thinking: Opus 5 thinks by default, bills the thinking as output, and never shows it to you. One request setting closes the gap to exact parity, and it is a setting the bigger Fable 5 refuses to accept. Opus 5 went GA on 2026-07-24, positioned as Fable-5-level intelligence at half the token price; whether your bill actually halves depends almost entirely on this one choice.
TL;DR
Default Opus 5 billed 3.1x the same-priced Opus 4.8 on our five-task matrix; 42-95% of its output tokens were hidden thinking.
thinking: {"type": "disabled"} brought Opus 5 to exact parity with 4.8 (384 vs 384 output tokens), accuracy held; Fable 5 rejects that parameter.
On agent traffic the tax collapses to +33%, with tool/batch scenarios near parity: adaptive thinking barely fires in tool loops.














