(Image credit: d-Matrix)

This Tom's Hardware Premium article is free to read with a Tom's Hardware account; no payment necessary. We're offering free access from August 23 to 26 so you can read all of our reporting from Hot Chips.d-Matrix presented Raptor, which it calls the first 3D DRAM accelerator for generative inference, at Hot Chips 2026 this week, showing a TSMC 4nm compute die bonded face-to-face at a 36-micron pitch on top of a custom-designed DRAM die that delivers 100 TB/s of bandwidth from 32GB per card.Co-founder and CTO Sudeep Bhoja put the vertical interface's energy cost at 0.37 pJ/bit against roughly 2.4 pJ/bit for moving data into an HBM4 base die, calling it "a measured number" from working silicon, and the accompanying ISCA 2026 paper, written with the University of British Columbia, projects around 4.7 times higher throughput per card than HBM-based designs. Bhoja, however, didn't disclose who manufactures the DRAM die.

(Image credit: d-Matrix)CEO Sid Sheth told CNBC in June that Raptor is slated to launch in 2027, but at Hot Chips the company gave no firm date or information on volume and pricing, and every performance figure shown, including 988 tokens per second per user on the 2.8-trillion-parameter Kimi K3 model at 1M-token context, is a d-Matrix projection built on early silicon.The custom DRAM dieRaptor inverts the usual 3D stacking arrangement by putting the logic die on top and the DRAM underneath, so a cold plate sits directly on the compute silicon and the DRAM die doubles as the interposer, carrying PCIe and die-to-die signals down through its TSVs. "For the same amount of power, you can drive the bandwidth up," Bhoja said during the session, "and so we were able to drive the bandwidth up here to 100 terabytes per second."The one-high stack achieves a power density of roughly 0.5W per square millimeter, which liquid cooling can handle, but the DRAM is designed for a junction temperature of 105°C, where retention collapses from a standard 32ms to 4ms, and the memory must refresh eight times more often. d-Matrix absorbed that penalty by shrinking each microbank to 1,366 rows and about 5.33MB, so a full refresh sweep costs only 1.37% of overall bandwidth.