Prashanthi Nuthi, VP at Enlace Health, helps organizations scale AI, technology and operations with strategic clarity and measurable impact.gettyOrganizations are moving quickly to embed AI into products, internal tools and business workflows. Most conversations understandably focus on what AI can do, like improve productivity, automate repetitive work and help teams move faster.But in focusing on the results, I've found that many leaders often skip some fundamental questions. For example, one question that is asked far less often is, "How will AI be managed once it becomes part of everyday operations?"That realization hit me after what seemed like a routine internal project. One of our quality engineers built a small utility using GitHub Copilot. The tool itself was fairly common, solved a real problem and worked exactly as intended, but the token consumption increased enough to trigger internal review.​What I realized after the fact was that, for a relatively straightforward task, an advanced model had been selected when a smaller, less expensive model would likely have produced the same result. Even though everyone ​made the right engineering decision, we hadn't optimized for cost awareness. We asked, "Which model is the most capable?" when we should have asked, "Which model is appropriate for this workload?" Those are very different questions.AI costs behave differently than other technologies.During cloud modernization programs, I became accustomed to infrastructure costs that were relatively predictable, as storage grows gradually, compute can usually be forecasted and capacity planning is something architecture teams understand well.With AI, every prompt, every summarization request, every customer interaction and every automated workflow consumes tokens. None of these individual requests may appear significant on their own, as the median price of 1 million input tokens is only $1, according to BenchLM. The challenge is what happens when hundreds of teams begin creating useful AI-powered experiences at the same time. For instance, at many companies today, they may have one team building a customer assistant, another adding document summarization and a third automating ticket classification.Each project is completely reasonable on its own, but they become a meaningful operational expense that few organizations have historically managed, leading to growing accounts of "AI sticker stock" across corporate America. Unlike adding employees or purchasing infrastructure, this growth often happens quietly. AI adoption expands one workflow at a time until leadership suddenly realizes that usage and spending has accelerated much faster than expected.Capability shouldn't be the default decision.One pattern I've noticed is that engineers naturally choose the most capable model available when guidance doesn't exist.​ Engineers are measured on whether a solution works, so that instinct makes sense.However, architecture leaders have a different responsibility. They have to ensure the solution continues working sustainably six months later.Those perspectives are complementary, but they aren't identical.In my experience, many internal workloads, including summarization, classification, extraction and structured content generation, don't always require the largest available models. In many cases, smaller models deliver results that are entirely acceptable while consuming significantly fewer tokens.The leadership conversation should shift from selecting the smartest model to selecting the most appropriate model.​Governance matters earlier than most organizations expect.One lesson I've taken from both cloud transformation and AI adoption is that governance is easiest to establish before usage accelerates. Once dozens of teams begin building independently, introducing standards becomes much harder. When organizations ask me where to begin, I encourage them to focus on visibility before optimization. Here's how: • Assign ownership for every production AI workload so someone is accountable for its cost and business outcome.• Establish simple dashboards that show which applications consume the most tokens and how usage changes over time.• Require teams to document why a particular model was selected instead of defaulting to the largest one available.• Review AI workloads periodically, just as infrastructure and security reviews are already part of engineering governance.• Measure business value alongside token consumption so discussions focus on return rather than cost alone.None of these practices slow innovation. In my experience, they actually accelerate it because teams make architectural decisions with better information.AI needs financial discipline, not slower adoption.I'm optimistic about AI. I've seen firsthand how quickly it can improve engineering productivity and unlock new possibilities for organizations.But I've also learned that successful AI adoption requires more than deploying powerful models. It also requires operational discipline.While developers will naturally optimize for capability and architects will optimize for scalability, leadership has to optimize for sustainability.To do this, leadership teams must start with conversations about where AI creates value, where it creates cost and how to intentionally balance both. As AI becomes part of everyday business operations, token economics must evolve from a technical metric to a key leadership responsibility.​​Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?