AI training clusters are reaching scales where physical infrastructure has become a critical design consideration. This whitepaper from AFL examines the physical architecture required to support a 16K-XPU AI training cluster, focusing on the back-end scale-out fabric, fiber density, optical connectivity and long-term scalability.
The paper presents a practical two-tier Clos reference architecture using commercially available optical connectivity technologies, alongside design principles intended to minimize future physical modifications as infrastructure evolves. It explores how backbone trunks, zone cabling and equipment cords can be structured to separate permanent infrastructure from topology-specific connectivity.
The reference design demonstrates how panel-to-panel connectivity, factory-terminated assemblies and topology implemented within equipment cords and zone cabling can create a physical network that is simpler to deploy and operate while remaining adaptable to future compute generations.






