If you thought 72 GPUs per rack was dense, next-generation designs from Nvidia and others will cram hundreds or even thousands of accelerators into a single massive system. But for any of that to happen, they're going to need a lot of optics and technology that's reminiscent of the old-fashioned telephone switchboard.This reality has fueled a flurry of investment in everything from photonics startups to established optical equipment and fiber manufacturing. In March, Nvidia invested $6 billion ($2 billion apiece) in Coherent, Lumentum, and Marvell to advance their optics tech.Some of these investments went to support optical circuit switching (OCS) technology. On Wednesday, Nvidia joined Maverick Silicon and Light Street Capital to spend another $125 million to support the development of iPronics' second-generation OCS tech.
Why switch packets when you can switch light?
Optical circuit switches are switches only in the literal sense. Unlike a Broadcom Tomahawk or Marvell Teralynx ASIC, an OCS appliance can't switch packets. The opto-electrical appliances are the modern equivalent of a telephone switchboard. But rather than human operators manually patching together two telephone lines, a high-speed actuator built using microscopic mirrors, piezoelectric actuators, LCDs, or other technologies reconfigures the network in the literal blink of an eye.OCS isn't particularly common in modern GPU deployments, but they have been used in AI clusters for quite a while. Specifically, Google has used OCSes in its TPU clusters for years now.Traditionally, Google's TPU pods have employed 2D and 3D torus topologies where accelerators communicate in a great big mesh, rather than relying on packet switched fabrics to connect them all together. The trade-off with mesh networks is potentially higher chip-to-chip latency and rigidity. On their own, they're not exactly the most flexible topologies out there. If you want to add, remove, or swap a dead accelerator, someone or something has to reconfigure the network.That something, in Google's case, is optical circuit switching. Traffic from Google's TPU clusters is transmitted optically through OCS appliances, which allows the Chocolate Factory to do things like change the pod size on demand or virtually hot-swap failed accelerators.Optical circuit switching is arguably the reason why Google is able to operate some of the largest single compute domains in the industry. Its network architecture isn't limited to packet switch radix.Smoothing over OCS' rougher edgesExisting OCS appliances aren't perfect, however. Many use micro-electromechanical systems (MEMS) devices to adjust which components are connected by moving microscopic mirrors. This approach works, but it's not what you would call fast. Reconfiguration times of about 100 ms are commonly quoted.OCS appliances also tend to be quite large, in part because of the mechanisms involved, but also because of the connectors used.







