Inference Is the New Oil- And Most of It Is Sitting Idle

Two-thirds of all AI compute is now inference. Not training. Not research. Serving requests, in production, right now.

That single number — flagged by hardware analyst Derek Colley when comparing Trainium3 and Nvidia's NVL72 rack specs — reframes the entire AI infrastructure conversation. Training built the models. Inference is where the money actually prints. And the industry is only just catching up to that reality.

The Shift Nobody Announced

For years, "AI compute" was shorthand for training runs: the multi-million-dollar, months-long jobs that produced GPT-4 or Claude. Infrastructure investment followed that assumption. Chip design followed it. Investor narratives followed it.