Industry
AMD and Cerebras Pair Up on Ultra-Fast AI Inference Hardware
16:00 · July 29, 2026
The partnership splits AI inference into two specialized stages rather than running the whole workload on one type of chip. AMD’s Helios rack-scale systems, built around its Instinct GPUs, handle prompt processing and large context windows, the compute-heavy front end of a request, while Cerebras’ Wafer-Scale Engine takes over token generation, the stage where extremely low latency matters most for applications like real-time chat or agentic workflows. Operating together as a single disaggregated inference pipeline, the companies say the combination delivers up to five times more tokens generated per second per watt compared with running the workload on Cerebras’ Wafer-Scale Engine alone, based on modeling the companies conducted this month. Cerebras plans to physically deploy AMD Helios systems inside its own data centers to support the joint architecture, and the combined offering is expected to become available initially through Cerebras Cloud in the second half of 2026. The announcement reflects a broader trend in AI infrastructure toward disaggregating inference workloads across specialized hardware rather than treating every stage of a request the same way, as providers look for ways to squeeze more throughput out of increasingly expensive and power-constrained data center capacity.