AI inference chip startup d-Matrix wants its upcoming Raptor XPU to live inside Nvidia’s MGX rack architecture by the end of 2027, and the company has a fairly concrete roadmap to get there. The Santa Clara-based firm expects its first Raptor tape-out to land before the close of 2026, with full rack-scale deployment targeting Q4 2027.

D-Matrix is integrating Raptor with Nvidia’s NVLink Fusion interconnect, which means up to 144 Raptor XPUs could operate inside a single NVLink fabric by the time 2027 wraps up.

What Raptor is actually built to do

Raptor debuted publicly at Hot Chips 2026 in August, and the technical specs are worth paying attention to. The platform combines a TSMC 4nm logic die with 3D-stacked DRAM, targeting over 100 TB/s of memory bandwidth with 32 GB of capacity per card.

D-Matrix claims Raptor delivers approximately 4.7 times higher throughput per card compared to HBM-based alternatives for generative inference tasks.