AI inference gets a new tier as context windows grow

AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they build longer contexts and generate more data that they must access quickly during inference.

Agentic AI is also changing the shape of the data problem. Interactions are growing longer and producing more information. At the same time, the data’s size, importance and movement through the infrastructure can all affect how quickly an application responds, according to Scott Shadley (pictured, left), director of technology planning at Solidigm Inc.

“One of the beautiful things that’s happened in this agentic AI, or even just the AI era, is [that] people are starting to pay attention to storage,” he said. “What’s unique about this particular era is it’s no longer one- or two-dimensional. We have data magnitude and growth in size, importance and all of the other volumetric aspects of that. But at the end of the day, it comes down to that bit of data and how fast that bit of data moves from point A to point B.”

Shadley, along with Anat Heilper (center), director of AI architecture at Vast Data Inc., and Ben Lee (right), director of solution management at Super Micro Computer Inc., spoke with theCUBE Research’s Rob Strechay during the Supermicro Open Storage Summit interview series. They discussed how the companies’ respective technologies fit together as expanding context windows and KV caches create new storage and memory demands for agentic AI.