Andrew Feldman, CEO of Cerebras Systems, announced a partnership with AMD on July 23 that pairs two very different chip architectures into a single inference platform. Cerebras shares climbed roughly 5% on the news.
The collaboration merges AMD’s Helios rack-scale solutions with Cerebras’ Wafer-Scale Engine, a chip literally the size of a dinner plate. Together, they form what the companies call a “disaggregated inference solution” designed for ultra-low-latency token generation.
What the partnership actually delivers
The joint system targets workloads that need both real-time responsiveness and raw throughput. Think agentic AI, coding assistants, and robotics, applications where a few hundred milliseconds of latency is the difference between useful and useless.
The companies project up to five times the tokens per second per watt compared to existing solutions.








