Cisco and Nvidia take AI factories from rack to runtime
AI factories are moving from ambitious plans toward production, but the path from graphics processing unit acquisition to usable systems remains a race against time.
Neoclouds already have customers waiting for capacity, enterprises are looking to bring inference workloads closer to home and sovereign AI programs are being built now. Those distinct buyer motions are converging around a common need: getting AI systems into production quickly enough to support the business, according to Will Eatherton (pictured, left), senior vice president of Cisco Systems Inc.
“Enterprises, many of them … are spending a large amount on tokens right now,” he said. “The rush and the pressure is getting these systems up so they can start offloading what has been an [application programming interface] into using local inference. I think it’s all converging on common architectures [and] common systems, but I think speed is either what’s broken or the challenge.”
Eatherton, along with Gilad Shainer (center), senior vice president of networking at Nvidia Corp., and Marc Hamilton (right), vice president of solutions architecture and engineering at Nvidia, spoke with theCUBE Research’s John Furrier at the “Cisco Secure AI Factory With Nvidia Expands to Rack Scale” event during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the rise of AI factories, rack-scale compute, Nvidia’s reference architecture, networking and the challenge of moving from the first token to sustained operations. (* Disclosure below.)









