AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production

The artificial intelligence infrastructure market is crossing an important threshold. The conversation is shifting from acquiring graphics processing units to building complete AI factories that can generate tokens reliably, efficiently and at scale.

This is the next bottleneck. GPUs may be the engine, but an AI factory is a system. Compute, networking, storage, cooling, software and operations must work together from Day 0 through continuous production. If one component fails to perform, expensive capacity sits idle and revenue slips away.

In a series of exclusive interviews on theCUBE, I spoke with Cisco’s Will Eatherton, senior vice president and head of networking engineering, and NVIDIA’s Gilad Shainer, SVP of networking, and Marc Hamilton, VP of solutions architecture and engineering, about how the companies are tackling that execution challenge as neoclouds, sovereign AI programs and enterprises move infrastructure into production. We explored the engineering, operational and financial requirements behind building AI factories at rack scale.

That is the strategic context behind the expansion of the Cisco Secure AI Factory with NVIDIA to rack-scale systems. Cisco Systems Inc., and NVIDIA Corp. are bringing together liquid-cooled compute, AI-optimized networking, validated designs and unified operations in an effort to compress the path between ordering infrastructure and producing the first token. The full rack-scale Secure AI Factory solution will be orderable through Cisco in September.