Cerebras and AMD partner to build the world’s fastest disaggregated AI inference solution

Disaggregated AI inference is proving to be more than a complementary answer to the prefill and decode bottleneck slowing enterprise AI at scale, and Cerebras and AMD just announced a partnership to build the fastest version of it in the world.

The recent collaboration pairs AMD’s Helios rack-scale architecture for the compute-intensive pre-fill phase with the Cerebras Wafer-Scale Engine for ultra-low-latency decode, according to Julie Choi (pictured), chief marketing officer at Cerebras. The resulting combination delivers 5x higher tokens per second per watt compared to existing solutions. Later this year, Cerebras will bring AMD Helios systems into its own data centers to power the pre-fill layer of the production deployment.

“Lisa and Andrew both announced how AMD and Cerebras are collaborating on the world’s most powerful disaggregated inference solution,” Choi said. “It’s a one plus one equals five X in this case.”

Choi spoke with theCUBE’s John Furrier at the Neo4j GraphTalk event during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the technical architecture behind disaggregated AI inference, along with what workloads are driving the fastest demand and why the partnership extends to deploying AMD Helios inside Cerebras data centers before the end of 2026.