Most of us pick a model the same way: read a leaderboard, pick the best one we can afford, ship it. Then the bill arrives and the "cheap" model turns out not to be the cheap one.

The reason is that there is no such thing as a cheap model. There is only a model that is cheap for the shape of your traffic — and the ranking reorders when the shape changes.

Quick vocabulary, because the whole argument lives in two words. A token is roughly ¾ of a word; models bill per million of them. Input tokens are what you send (prompt, files, chat history); output tokens are what the model writes back. They have different prices, and the gap between them is not the same for every model.

The number you don't have yet

Every model's price is two numbers, and every vendor publishes the ratio between them without commenting on it. Here it is, list prices per 1M tokens, snapshot taken 12 Aug 2026: