Over the past two years, many enterprises have taken a maximalist approach to AI adoption — launching more use cases, onboarding more employees, deploying more agents, and consuming more tokens. For a new technology, this was rational. Enterprises did not yet know which workflows would be transformed, which models were reliable enough or which employees would adopt the tool. The fastest way to find out was to deploy as many agents and run as many pilots as possible to figure out what worked.

But this year, the conversation changed. In nearly every executive review, the question is no longer “How do we get more AI?” or “Which model is smartest?” It’s “How do I generate impactful outcomes, while keeping AI costs under control?” Different companies have different answers. Uber tightened internal AI spending after consumption ran far ahead of plan. Anthropic, OpenAI and Google have aggressively cut input token prices and rolled out prompt caching. And Cursor, with Composer 2.5, now considers cost an important factor in model selection — not just performance.

Cutting token and model costs are one solution, but they ignore the more profound issue: Most companies are not architected to use tokens efficiently. The modern AI pipeline behaves more like a sieve — leaking tokens, and dollars, at every phase of execution. Pouring cheaper tokens through a leaky foundation isn’t a sustainable solution. To truly generate great, cost-effective outcomes, companies need an architectural foundation that closes those leaks. Here’s how.