Last month I sat watching one of my own agent pipelines grind through a step that a human was actively waiting on. The job itself was fine. The model was fine. The problem was that every interactive request in the system, from "summarize this PDF right now because a user is staring at a spinner" to "re-embed this archive of old documents overnight," was treated as exactly the same kind of work. Same model, same tier, same speed. That is a design flaw, and this week OpenAI turned it into a line item.
On July 30, OpenAI restructured GPT-5.6 pricing along two axes at once. Luna, the smallest tier, dropped 80 percent. Terra dropped 20 percent. And Sol, the flagship, gained something new: Fast mode, up to 2.5 times the speed of Standard processing, at twice the price, with what OpenAI explicitly states is no change in intelligence.
That last combination is the one worth thinking about. A smarter model is a capability decision. A faster model at 2x cost is a pure money-for-latency trade, and most teams have never written down the rules for when that trade makes sense. So I did the thing I do before spending money: I collected the real numbers, ran the arithmetic on a workload like mine, and built a decision matrix. It is below, and you can steal it.






