Small Models Have Arrived — And They Change the Economics of Everything
GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens on OpenRouter. With prompt caching, that drops to $0.02 per million cached input tokens. At that price, a complex multi-step agent workflow that used to cost $1 is now $0.10.
This is not a marginal improvement. It's a phase transition in what you can build.
The HN thread on Calvin French-Owen's "Small Models Have Arrived" (469 points, 39 comments) captured the sentiment concisely: "Those of us without fable-sized expense accounts noticed this quite a while back." But the data bears out that the inflection point is here, across every major model provider.
Let me show you the numbers, the architecture changes that enabled them, and what they mean for production engineering.






