Everyone's watching the AI price war for the wrong reason. The headlines are about how cheap tokens got. The actual story is what cheap tokens do to how you should be architecting software right now.
Here's what changed, and why it matters more than the price cut itself.
Models aren't one thing anymore, they're tiers
The major labs have quietly split their lineups into tiers. Cheap, fast models for routine work. Expensive, deep-reasoning models for the hard stuff. This isn't a pricing gimmick, it's an architecture signal. If your app sends every request to the same model regardless of difficulty, you're either overpaying for simple tasks or underpowering the hard ones.
The pattern worth adopting: route by task, not by app.







