TL;DRNvidia confirmed Vera Rubin is in full production, with OpenAI deploying at scale in Q3 and CoreWeave reporting ten times the token output.
Nvidia confirmed on Monday that its Vera Rubin platform has reached full production, with Ian Buck, the company’s vice president of accelerated computing, telling reporters at Nvidia headquarters that systems are now shipping to customers including OpenAI, CoreWeave, Google Cloud, Microsoft Azure, Meta, and Dell. OpenAI plans to adopt Vera Rubin at scale during the third quarter, according to Bloomberg, which first reported the briefing.
CoreWeave, one of the first cloud providers to receive the hardware, told Bloomberg that its NVL72 racks are delivering ten times the token output of the previous generation. The NVL72 is a full-rack system that pairs 72 Rubin GPUs with Vera CPUs and uses liquid cooling to eliminate internal cabling, a design change Nvidia demonstrated at the headquarters event. Buck said the cooling approach allows the company to remove copper connections that previously limited how tightly components could be packed together.
The Vera CPU is the piece Nvidia built to replace the processors it previously bought from others, and the company used the briefing to draw direct comparisons with AMD. Nvidia claimed the Vera CPU is nearly twice as fast as AMD’s Turin chip on Python workloads, a benchmark chosen because Python dominates the software stack that runs most AI inference. Anthropic and OpenAI are among the first labs to receive the processor, alongside Perplexity, SpaceX, and Oracle.











