Nvidia is weighing a significant spec change to its upcoming Rubin Ultra GPU, potentially shipping the chip with far less high-bandwidth memory than the company originally targeted. The adjustment, reported by TrendForce in early August 2026, is a direct response to tightening supply of HBM from the three major producers: SK Hynix, Samsung, and Micron.
The original vision for Rubin Ultra was ambitious. When Nvidia first unveiled the chip at its GTC developer conference in 2025, configurations on the table included up to 1 TB of HBM4E memory. The current Vera Rubin GPU, already in mass production, ships with up to 288 GB of HBM4 in 12-Hi stacks, delivering roughly 22 TB/s of memory bandwidth. Rubin Ultra was supposed to push well past that.
Now, Nvidia is testing at least three variants with meaningfully reduced memory. One configuration under review uses 8-Hi HBM4 stacks carrying approximately 192 GB of memory. Compared to the 288 GB on the current Vera Rubin, that is a 33% reduction in on-chip memory for a chip that was supposed to represent a generational leap forward.
Why memory is everything for AI workloads
High-bandwidth memory is not just about capacity. The “bandwidth” part refers to how fast data moves between the memory and the chip’s processing cores. At 22 TB/s, the current Vera Rubin already operates at speeds that dwarf conventional server memory by orders of magnitude.











