NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two parts. Time-to-rack runs from silicon leaving the fab to an assembled system arriving on a data center floor. Time-to-token covers everything thereafter: power, cooling, networking, and the software stack that makes the infrastructure productive on day one.
NVIDIA Grace Blackwell NVL72 platforms draw on millions of parts and thousands of suppliers spread across the globe, and the final system is built by dozens of OEMs and ODMs. Just one compute tray–one of eighteen in a single rack–requires two NVIDIA Grace CPUs, four NVIDIA Blackwell GPUs, and thirty-two HBM3e stacks. The supply chain we created for Vera Rubin is twice as large as Grace Blackwell. CPUs, GPUs, and memory are all critical components, and the availability of each changes from week to week, so the part holding up a build one week may be freely available the next. Each carries its own bill of materials, its own suppliers, and its own lead times. Multiply that by every sub-assembly in the rack, and the result begins to look less like a supply chain and more like a daunting combinatorics problem.







