Token per watt becomes the defining metric as storage moves to AI’s critical path
Token per watt — not raw compute — is emerging as the defining efficiency metric for AI data centers, putting storage at the center of an infrastructure rethink that is reshaping how the industry measures performance, cost and scale.
As agentic AI drives an explosion in context memory demand, the role of solid-state storage has shifted from afterthought to critical path. KV cache offload, inference data pipelining and expanding context windows are converging to create a new storage tier inside AI clusters — one that determines whether GPUs stay busy or sit idle, according to Avi Shetty (pictured), vice president of AI ecosystem, solutions and market enablement at Solidigm, a trademark of SK Hynix NAND Product Solutions Corp.
“The new metric for data centers is token per watt, for dollar per token per watt,” Shetty said. “A single prompt — like, we are in Paris, hey, top five recommendations for restaurants — that simple prompt translates to around 13 gigabytes of data when you account for system activities, tool activities, database requests. When an eviction occurs, you need to rebuild the cache. When you rebuild the cache, your GPU is waiting — and that cannot happen.”








