No one likes a surprise sky-high bill. But that’s exactly what many companies have faced this year as they deploy popular AI coding agents like Claude Code and Codex for increasingly long-running, autonomous tasks. Unlike a chatbot conversation, these agents can work for hours, repeatedly calling frontier models and quietly racking up millions of tokens. A developer can go to lunch and return to find that his agent has spent thousands on the inference, or output of the model. The result? Sticker shock.
A new study illustrates the impact: 62% of organizations said an unexpected AI expense materially altered a business decision over the past year. Among them, 40% said the issue required board-level escalation, 33% implemented emergency spending freezes, and 25% delayed or canceled an AI initiative outright.
That’s why AI model routers, software that allows organizations to choose the AI model for the right task at the right cost, have suddenly become one of the hottest areas in enterprise tech. Rather than sending every request to the most expensive frontier model, companies can either manually define or automatically choose the model that offers the best combination of cost, speed, and performance for each step of an agent’s work. Companies say this kind of intelligent model routing can reduce inference costs by double-digit percentages – in some cases up to 30%.









