TDK — leaderboard ()

Disaggregated AI inference pairs AMD Helios with Cerebras WSE, due H2 2026

AMD and Cerebras are collaborating on a disaggregated AI inference solution that pairs AMD Helios rackscale systems with the Cerebras Wafer-Scale Engine (WSE) in a single inference workflow. The aim is to combine a high-throughput prompt stage with low-latency decode and token generation for latency-sensitive inference use cases.

In the design described, AMD Helios handles prompt processing and large context windows, while Cerebras WSE accelerates the memory-bandwidth-intensive token-generation phase. The companies describe the combined workflow as a way to tune the two primary stages of inference independently, rather than forcing one architecture to handle both.

AMD and Cerebras expect the joint configuration to deliver up to 5x higher tokens per second per watt. That figure is based on modeling done by AMD Performance Labs and Cerebras in July 2026, comparing an AMD Helios rackscale solution plus Cerebras WSE against a Cerebras WSE-only configuration at a “comparable interactivity point” using the Kimi 2.6 1T Model. It’s also expressed as tokens per second per kilowatt (TPS/kW), and the companies note that system configurations may vary by manufacturer.

For data center operators, the disaggregated approach is the practical takeaway: it implies two different compute domains working as one inference pipeline, which can shift how clusters are sized and where power is spent. If the workflow split holds up in real deployments, it could let operators scale the prompt stage and the token-generation stage independently to match application needs, especially where response time is the constraint.

Cerebras plans to deploy AMD Helios systems in its data centers. The joint solution is expected to be available first through Cerebras Cloud in the second half of 2026.

AMD CEO Dr. Lisa Su said, “AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach.” Cerebras CEO and co-founder Andrew Feldman said, “Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers.”

AMD Helios was unveiled with the Cerebras WSE workflow at Advancing AI 2026, alongside AMD Instinct GPUs referenced as part of the throughput side of the design.

Source: AMD

Get Data Center Engineering News In Your Inbox:

By subscribing, you agree to our Privacy Policy and Terms of Use. You can unsubscribe at any time.

Popular Posts:

Johnson-Controls-has-launched-an-Absorption-Chiller-Reference-Design-Guide-aimed-at-data-centers-running-on-on-site-power-generation
Johnson Controls' new absorption chiller guide targets 44% lower cooling power in AI data centers
CPC-TRM-Tech-Guide-Liquid-Cooling-Reliability-(1)-4
In liquid cooling reliability, tubing choice creates another layer of protection
Why data center power validation is moving to real-time simulation
Why data center power validation is moving to real-time simulation
trm-DC4-landing-page-banner
CPC says its full-flow AI cooling connectors cut pressure drop up to 90%
Warm-Water-Ready-Accelsius-Whitepaper-3
Accelsius two-phase cooling runs NVIDIA B200 9° C cooler than single-phase
TDK — tower ()

Share Your Data Center Engineering News

Do you have a new product announcement, webinar, whitepaper, or article topic? 

Get Data Center Engineering News In Your Inbox:

By subscribing, you agree to our Privacy Policy and Terms of Use. You can unsubscribe at any time.