The AI hardware conversation is stuck on one word: compute. A startup out of Tel Aviv wants to change the word to memory. Majestic Labs, founded in 2023 by former Google and Meta engineers, has unveiled a server it says can do the work of a rack of Nvidia GPUs. It does so by attacking a different bottleneck.
The pitch, reported by TechRadar, is that pairing pricey GPUs with scarce high-bandwidth memory has become a dead end for AI inference. Running a model is often limited by how much fast memory you can reach, not raw compute. So Majestic ditched the GPU.
Its server, Prometheus, swaps graphics chips for what it calls Ignite AI Processing Units. Each unit blends Arm cores with RISC-V vector and tensor engines. Up to 12 sit in one server, sharing a single pool of 8TB to 128TB of LPDDR6. That is the cheap memory found in phones, not the costly high-bandwidth memory that GPUs depend on.
A different way to hit the memory wall
The trick is the wiring. Instead of bolting memory onto each GPU package, Majestic pools it through custom aggregation chiplets. Copper cables up to a metre long connect them. The result, it claims, is one coherent pool far larger than a GPU box can address.








