Skip to content

AMD and Cerebras Pair Helios GPUs with Wafer-Scale Engine for Low-Latency AI Inference

AMD and Cerebras will pair AMD Helios GPU racks with the Cerebras Wafer-Scale Engine for low-latency AI inference, arriving via Cerebras Cloud in H2 2026.

Two speakers onstage at AMD Advancing AI beside a slide charting up to 5x throughput per kilowatt for Helios plus Cerebras
Credit: AMD

AMD and Cerebras Systems announced a technical partnership on July 23 that pairs AMD’s Helios rackscale systems with the Cerebras Wafer-Scale Engine in a single disaggregated AI inference workflow, the companies said at AMD’s Advancing AI 2026 event in San Francisco.

The split is deliberate. AMD Helios handles the throughput-heavy side of an inference request, processing prompts and large context windows across AMD Instinct GPUs, while the Cerebras Wafer-Scale Engine (WSE) takes over the memory-bandwidth-intensive decode stage, generating tokens with what Cerebras calls ultra-low latency. AMD and Cerebras describe this as disaggregated inference: routing the two stages of a request to whichever hardware suits it, rather than running an entire workload on one type of chip.

The headline number in both AMD’s and Cerebras’ releases is a claim of up to 5x higher tokens per second per watt from the combined system. That figure is not a comparison against a competitor’s hardware: per a footnote attached to both releases, it comes from July 2026 modeling by AMD Performance Labs and Cerebras that measured tokens per second per kilowatt at a comparable interactivity point running the Kimi 2.6 1T model, pitting an AMD Helios-plus-Cerebras-WSE configuration against a Cerebras WSE-only setup. Cerebras’ own site adds a standing caveat that its performance comparisons “are based on third-party benchmarking or internal testing” and that results “may vary depending on workload, configuration, date and models being tested.”

AMD chair and CEO Dr. Lisa Su said the arrangement extends the Helios lineup “into the most latency-sensitive applications,” positioning it as a base for “real-time agentic AI.” Cerebras co-founder and CEO Andrew Feldman said pairing the Wafer-Scale Engine with AMD’s rack-scale systems gives Cerebras “an incredible opportunity to bring that performance to even more customers,” according to the companies’ joint statements.

What is shipping first is a deployment plan, not a retail product. Cerebras plans to install AMD Helios systems in its own data centers and run them alongside its Wafer-Scale Engine, with the combined setup reaching customers first through Cerebras Cloud in the second half of 2026, per both companies. Neither side published independent benchmark results, a customer list, or pricing alongside the announcement, so the 5x efficiency claim rests on the companies’ own modeling until the joint system is running in production and someone outside AMD or Cerebras can test it.

Share this story

Thomas Albrecht

Thomas Albrecht covers processors and graphics silicon for techshooked, from desktop CPUs to the accelerators behind modern AI. His standard is numbers-first: document the test conditions, report the results a spec sheet leaves out, and judge a chip on measured performance rather than launch slides.