AMD and Cerebras are collaborating on a disaggregated AI inference solution that pairs AMD Helios rackscale systems with the Cerebras Wafer-Scale Engine (WSE) in a single inference workflow. The aim is to combine a high-throughput prompt stage with low-latency decode and token generation for latency-sensitive inference use cases.
In the design described, AMD Helios handles prompt processing and large context windows, while Cerebras WSE accelerates the memory-bandwidth-intensive token-generation phase. The companies describe the combined workflow as a way to tune the two primary stages of inference independently, rather than forcing one architecture to handle both.
AMD and Cerebras expect the joint configuration to deliver up to 5x higher tokens per second per watt. That figure is based on modeling done by AMD Performance Labs and Cerebras in July 2026, comparing an AMD Helios rackscale solution plus Cerebras WSE against a Cerebras WSE-only configuration at a “comparable interactivity point” using the Kimi 2.6 1T Model. It’s also expressed as tokens per second per kilowatt (TPS/kW), and the companies note that system configurations may vary by manufacturer.
For data center operators, the disaggregated approach is the practical takeaway: it implies two different compute domains working as one inference pipeline, which can shift how clusters are sized and where power is spent. If the workflow split holds up in real deployments, it could let operators scale the prompt stage and the token-generation stage independently to match application needs, especially where response time is the constraint.
Cerebras plans to deploy AMD Helios systems in its data centers. The joint solution is expected to be available first through Cerebras Cloud in the second half of 2026.
AMD CEO Dr. Lisa Su said, “AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach.” Cerebras CEO and co-founder Andrew Feldman said, “Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers.”
AMD Helios was unveiled with the Cerebras WSE workflow at Advancing AI 2026, alongside AMD Instinct GPUs referenced as part of the throughput side of the design.
Source: AMD















