Michael Eichsteadt is VP of Engineering at iManage, a company dedicated to Making Knowledge Work.getty​Artificial intelligence has moved from experiment to established infrastructure across industries, and with that shift has come a new line item on the balance sheet. Yet the economics are misread. The price of any fixed level of AI capability has collapsed. Last year’s frontier performance now costs a fraction of what it did even as each new frontier model commands a premium over the one before.​At the same time, consumption is compounding with agentic workflows multiplying the computation a single task consumes. Intelligence is getting cheaper per unit, while enterprises are buying more expensive units and more of them. For technology buyers, this reality makes vendor evaluation more important than ever. The question that matters most becomes: What is a vendor doing about the economics underneath its product? The Cost Of Compute Every AI interaction carries a price. When a user submits a query, generates a document or triggers an automated workflow, that activity requires inference—the computational work of running a model and returning a result. Inference isn't free. It's measured in tokens, priced per million and billed to someone. Across the market, AI pricing does not yet reflect underlying cost. As AI utilization scales—particularly with the rise of agentic AI—the bill will eventually come due. Costs absorbed today are passed along tomorrow. Attention is turning to what actually drives the bill. The model sets the rate, the architecture sets the volume and the two multiply. Choosing the wrong class of model for a task or mismanaging context can swing costs by orders of magnitude. An agentic workflow that pushes entire document sets into each request will cost dramatically more than one that retrieves only what the task requires, regardless of the underlying model. Both disciplines can be engineered, and they are among the clearest signals that a vendor has done the work. Questions Every Buyer Should Be Asking Given this trajectory, technology buyers need to treat AI cost management as a core criterion in vendor evaluations—on par with security, scalability and reliability. The right questions to ask any vendor include: What steps are you taking to control inference costs? And how will those costs be passed along over time? Vendors that can demonstrate that they're actively working to reduce costs—not just pass them through—will stand out. Reselling access to LLMs is a legitimate way to bring capability to market quickly. It just offers no structural cost advantage on its own. Over time, the difference will show up in what a vendor has built around and beneath the model.The Case For Small Language Models One of the most promising strategies for managing AI costs is adopting small language models (SLMs). Unlike LLMs—which are trained on massive, general-purpose datasets—SLMs are compact models, many of which are open-weight, that can be fine-tuned on domain-specific data and deployed on privately owned or controlled infrastructure. This matters for two reasons. First, the cost difference is dramatic. Second, for specific, repeatable tasks, SLMs can match or outperform their larger counterparts. They excel at high-volume, repeatable work: extracting names and clauses, classifying documents and other targeted operations where precision matters. Note, however, that smaller isn’t always better. Frontier LLMs still earn their premium on open-ended reasoning, novel problems and low-volume work where flexibility matters more than unit cost. Also, the SLM landscape is evolving quickly: techniques like model distillation—using larger "parent" models to train smaller, highly efficient "child" models—are helping to push down cost and size. Capabilities that debut at the frontier reliably arrive in smaller models within months. The discipline to look for is portfolio thinking. Match the model to the task, and reserve expensive general-purpose inference for problems that need it. A vendor should be able to explain which tasks it routes to which model class and whether its cost-reduction strategy genuinely brings inference closer to home or simply repackages the same large-model costs in a new wrapper. Cost And Security Are The Same Conversation Cost optimization and security are deeply intertwined. When AI inference runs through a third-party LLM provider, data crosses into external infrastructure—meaning privileged or sensitive information is processed outside the buyer's direct control. No architecture keeps data entirely away from a model. Whatever context a model needs to do its job must be provided. The meaningful difference is governance. When data stays in a governed repository and models access it through standards like the Model Context Protocol, that access is permissioned, auditable and scoped to the task at hand, rather than being bulk-copied into a vendor’s intermediate systems, where visibility and control erode with every additional hop. Choosing Partners Who Can Keep Pace The economic pressure to control AI costs will only intensify as usage compounds. The vendors worth partnering with are those who treat this concern as an active area of investment. What matters is not just what a vendor is doing today but its willingness to stay nimble as the technology evolves. A vendor that cannot explain how it matches models to tasks, or one committed to a single class of model for every job, deserves a closer look. The best partners are already asking the same cost and performance questions you are, and building their solutions accordingly. The price of yesterday’s intelligence will keep collapsing, the frontier will keep charging a premium, and the bill will keep growing. The vendors worth betting on are the ones building for all three facts at once.Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?