Power is AI infrastructure’s inescapable constraint. How many tokens an AI factory can generate within a fixed power budget determines its revenue and profitability. Because of this, performance per watt — a metric that can’t be gamed, only earned through real-world results — is the foundation for AI factories.

As agentic AI drives token demand higher, the infrastructure decisions organizations make today will determine who scales and who doesn’t in a power-constrained world.

Virtually every frontier AI model today runs on a mixture-of-experts (MoE) architecture. Serving MoE at rack scale demands codesign across every layer of the system and software stack, plus the operational depth earned from running these models under real production load. With the NVIDIA Blackwell NVL72 platform, that rack-scale foundation is already built and proven, delivering the highest performance per watt to maximize revenues and the lowest token cost to maximize profit margins. It’s this foundation that the NVIDIA Vera Rubin platform builds upon next to further elevate rack-scale energy efficiency.

Maximizing Performance per Watt for Frontier AI

Each new generation of frontier models brings architectural changes that unlock greater intelligence while demanding new optimizations to run efficiently at scale.