AMD and Cerebras Systems announced a technical partnership on July 23 that pairs AMD’s Helios rackscale systems with the Cerebras Wafer-Scale Engine in a single disaggregated AI inference workflow, the companies said at AMD’s Advancing AI 2026 event in San Francisco.
The split is deliberate. AMD Helios handles the throughput-heavy side of an inference request, processing prompts and large context windows across AMD Instinct GPUs, while the Cerebras Wafer-Scale Engine (WSE) takes over the memory-bandwidth-intensive decode stage, generating tokens with what Cerebras calls ultra-low latency. AMD and Cerebras describe this as disaggregated inference: routing the two stages of a request to whichever hardware suits it, rather than running an entire workload on one type of chip.
The headline number in both AMD’s and Cerebras’ releases is a claim of up to 5x higher tokens per second per watt from the combined system. That figure is not a comparison against a competitor’s hardware: per a footnote attached to both releases, it comes from July 2026 modeling by AMD Performance Labs and Cerebras that measured tokens per second per kilowatt at a comparable interactivity point running the Kimi 2.6 1T model, pitting an AMD Helios-plus-Cerebras-WSE configuration against a Cerebras WSE-only setup. Cerebras’ own site adds a standing caveat that its performance comparisons “are based on third-party benchmarking or internal testing” and that results “may vary depending on workload, configuration, date and models being tested.”
AMD chair and CEO Dr. Lisa Su said the arrangement extends the Helios lineup “into the most latency-sensitive applications,” positioning it as a base for “real-time agentic AI.” Cerebras co-founder and CEO Andrew Feldman said pairing the Wafer-Scale Engine with AMD’s rack-scale systems gives Cerebras “an incredible opportunity to bring that performance to even more customers,” according to the companies’ joint statements.
What is shipping first is a deployment plan, not a retail product. Cerebras plans to install AMD Helios systems in its own data centers and run them alongside its Wafer-Scale Engine, with the combined setup reaching customers first through Cerebras Cloud in the second half of 2026, per both companies. Neither side published independent benchmark results, a customer list, or pricing alongside the announcement, so the 5x efficiency claim rests on the companies’ own modeling until the joint system is running in production and someone outside AMD or Cerebras can test it.













