Products
AMD and Cerebras pair up on ultra low latency AI inference architecture
9:00 AM · July 28, 2026
AMD and Cerebras Systems announced a technical partnership on July 22, 2026 at AMD's Advancing AI 2026 conference to build a disaggregated AI inference solution combining AMD's Helios rack scale systems with Cerebras's Wafer Scale Engine chips. In the joint design, Helios handles the high throughput work of processing initial prompts and large context windows, while Cerebras's specialized architecture takes over the memory bandwidth intensive job of generating tokens with ultra low latency, an approach the companies say can deliver up to five times more tokens per second per watt than running on Cerebras hardware alone. AMD CEO Lisa Su said during her keynote that “nobody predicted AI agents would grow this fast,” framing the partnership as a response to real time latency demands from autonomous coding agents, robotics and scientific discovery applications rather than traditional chatbot style workloads. Cerebras CEO Andrew Feldman said pairing with AMD “gives us an incredible opportunity to bring that performance to even more customers.” The joint solution is expected to become available first through Cerebras Cloud in the second half of 2026, with Cerebras also planning to deploy AMD Helios systems directly in its own data centers. The deal extends AMD's push to build a credible inference alternative to Nvidia's dominant position, following AMD's separate multibillion dollar chip and investment agreement with Anthropic earlier this month.