The AI inference race moves beyond GPUs to reshape data center infrastructure

AI inference infrastructure is becoming a system-level challenge as organizations move generative and agentic applications into production. Graphics processing unit performance remains essential, but storage latency, network bandwidth, data movement and power consumption increasingly determine the cost and speed of producing tokens.

The requirements also vary by workload. Interactive chat prioritizes latency, batch inference emphasizes throughput and agentic systems create expanding contexts. At the same time, retrieval-augmented generation and multi-tenant AI factories intensify demands across the stack, according to Ka Wai Leung (pictured, center), AI solutions product management at IBM Corp.

“You need to understand what type of workload,” he said. “Based upon the workload, you understand the characteristics of the workload, and you build your system behind it. That’s how you scale.”

Leung; William Li (left), general manager of solution management at Super Micro Computer Inc.; and Anders Graham (right), senior director of SSD marketing and business development at Kioxia Holdings Corp., spoke with theCUBE’s Rob Strechay for the Supermicro Open Storage Summit interview series, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed how workload requirements, data architectures and coordinated infrastructure design shape inference performance at scale. (* Disclosure below.)