Andrew Feldman, CEO of Cerebras Systems, announced a partnership with AMD on July 23 that pairs two very different chip architectures into a single inference platform. Cerebras shares climbed roughly 5% on the news.

The collaboration merges AMD’s Helios rack-scale solutions with Cerebras’ Wafer-Scale Engine, a chip literally the size of a dinner plate. Together, they form what the companies call a “disaggregated inference solution” designed for ultra-low-latency token generation.

What the partnership actually delivers

The joint system targets workloads that need both real-time responsiveness and raw throughput. Think agentic AI, coding assistants, and robotics, applications where a few hundred milliseconds of latency is the difference between useful and useless.

The companies project up to five times the tokens per second per watt compared to existing solutions.