Steven Carlini, Chief Advocate, Data Centers and AI, Energy Management Business Unit, Schneider Electric.getty​​Tokens have several definitions. First, they are the language native AI speaks because AI models are trained using data that is tokenized (transformed into tokens for processing). Working AI models use prompts to query the model, usually in the form of words broken into tokens. Responses from the model (predictions, content and reasoning) are generated as tokens during inference then transformed into words, images, songs, etc.Secondly, tokens are the primary AI currency. AI tokens have evolved into the fundamental unit of value. Cloud providers and model developers (such as OpenAI and Anthropic) primarily charge by the number of tokens—both output and input tokens.But token use is costing companies more than they expected. Budgeting for it is proving extremely challenging as companies ramp up maturity. This has created the perfect opportunity to go one step beyond the metric tokens per watt and make tokens per watt per dollar the new AI productivity metric.​Token Volume Not Necessarily Equal To Value "Tokenmaxxing" is a Silicon Valley workplace trend where employees maximize their AI token consumption to demonstrate productivity. Tech firms initially encouraged this practice, but because AI services bill and measure activity by tokens, costs were excessive—sometimes, more than the employee’s salary. Nvidia top engineers may use over $250,000 per year in tokens, and Meta is starting to cap tokens used by employees due to escalating costs.The issue is trying to understand if tokens produced by AI are doing constructive work. Sure, many tokens add value, but using a metric that only quantifies volume is not useful and may have negative repercussions. It’s like grading call center reps by the number of calls they take rather than how many problems they solve, perhaps incentivizing them to shorten calls without problem resolution.​Tokens Per Watt And IT Work Output To Power UsedA performance and power efficiency metric attempting to assign IT work output to power is valuable for benchmarking and spotting waste or AI model overkill. Do we need the largest frontier models running the latest accelerated compute to perform simple tasks like generating an email? It’s like driving a Ferrari to take your kids to school.To increase tokens per watt, a viable approach is to lower the power used. One way is adopting liquid cooling for AI inference hardware to reduce the energy used. Heat rejection in liquid cooling offers choices of power efficiency and water utilization that the user can design based on preferences. AI job scheduling to process intensive workloads when lower cost power is available will lower power use. Using the most efficient UPS and power distribution provides another way to lower power costs. Reference designs and prefabricated IT PODs that have validated performance specifications can ensure the desired power use. In addition, with digital twins on the power systems and cooling systems in AI factories, simulations can be run swapping out components and using different architectures to optimize power use.​A Better Metric: Tokens Per Watt Per DollarAs AI firms often charge by the token, better token efficiency yields lower cost. By cost I mean the total cost of the AI compute including the capital expenditure of purchasing or renting the hardware with the operational costs of electricity. Tokens cost money going in and out of AI models running inference. Processing queries is less expensive than generating new responses. Output tokens typically cost three to five times more than input tokens. Plus, pricing per million tokens fluctuates widely across frontier models. For example, high-tier models can charge anywhere from $5 million to $50 million output tokens, while smaller, verticalized and optimized models cost a fraction of a dollar.​Tokens per watt per dollar shifts the focus away from buying the most expensive hardware for the fastest AI model. It rewards models that are optimized to produce intelligence with optimized energy and capital costs.​All Tokens Not Created EqualOutput tokens must add value to positively affect business impact. But the model does not specifically account for quality that the tokens add. Plus, games could technically be played with the data processing. For example, FP8 (floating point 8-bit) allows for higher throughput compared to FP16 (floating point 16-bit), meaning more tokens can be processed per second per user. FP4 (floating point 4-bit) yields even higher numbers of tokens, but the accuracy and impact of the tokens could be less.​Using Tokens Per Watt Per DollarCompanies that measure tokens per wt per dollar could use it in many ways.1. Data center capacity planning is a prime use case. Data centers are increasingly becoming power-constrained as GPU power consumption increases. Companies cannot simply add new, higher token output GPUs as that increases power consumption significantly.2. Developers can use tokens per watt per dollar software to decide which AI model to deploy.Instead of using the most expensive frontier model, companies could align the model's capabilities with its energy and dollar cost.3. The model can be used as a tool to help track employee and AI agent ROI. AI agent expenses are reaching unexpected thresholds and businesses now need to monitor token usage at the individual employee and department levels. Companies can use this metric to validate the ROI of their AI implementations, seeing if the work output justifies the token and energy costs.​Tokens per watt per dollar is a smart framework for bringing energy costs into the way we track intelligence. At the micro level, it can help companies decide the best AI strategy to deploy. At the macro level, it could be a useful metric for understanding output in an information economy.Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?