- AMD and Cerebras will split inference across two machines, with Helios racks handling fast processing and the Wafer-Scale engine generating tokens, available through Cerebras Cloud in the second half of 2026.
- Nvidia is also doing something similar by licensing SRAM decoding technology from Groq, an AI chip startup, for $20 billion.
- The move sees AMD and Cerebras claim 5x higher tokens per watt compared to a standalone Cerebras WSE setup.
AMD and Cerebras Systems have announced a technical partnership that combines the former’s Helios rackscale system with the latter’s Wafer-Scale Engine in what both companies call a disaggregated inference solution.
The move has enabled a combined AMD Helios and Cerebras WSE configuration to deliver up to five times the tokens per second per watt (TPS/W) in internal tests conducted by both chip designers.
The move aims to address a Cerebras WSE efficiency challenge by offloading fast processing to AMD’s rack-scale offering.
Latest videos ofTechnologyRadar
A play focused on efficiency gains?
Both AMD and Cerebras Systems are presenting the news as a victory, and it very well could be, given the latter’s efficiency gains at stake and the former’s ability to gain access to SRAM decoding technology without spending the $20 billion that Nvidia shelled out late last year for a non-exclusive deal.
However, it should be noted that the 5 tokens per second per watt efficiency claims are compared to an existing Cerebras WSE (Wafer-Scale Engine) as a benchmark, while running the open source Kimi 2.6 1T model, making them impressive, but without a direct comparison to the numbers from a rack-scale Nvidia offering, one that lacks context, especially when efficiency is the metric.
The idea itself is solid and well established in the industry, and WSE is known to struggle with the ‘prefill’ part of the equation while handling the ‘decoding’ segment relatively well, essentially substituting AMD hardware where Cerebras equipment falls short.
The choice of the Kimi 2.6, however, deserves a second look. The Moonshot AI model, released on April 20, 2026, is an expert combination design with one trillion total parameters but only 32 billion assets per token, and ships natively on INT4. In INT4, the total weight set reaches approximately 500 GB. A single Cerebras wafer holds 44 GB. Even before the KV cache, a Cerebras-only implementation needs just over a dozen wafers just to hold the model, while a Helios rack could hold it about sixty times.
That asymmetry means that the five-fold figure is measured in a model that is close to the least favorable for a WSE-only configuration. A dense model small enough to rest on a handful of wafers could benefit Cerebras much more. None of this makes the number wrong, but it does justify the need for additional testing to demonstrate both its strengths and weaknesses for different models.
An association without numbers, for now
More importantly, the absence of financial information could well be a future story, especially at a time when there are growing concerns about “circular financing” in an industry where Nvidia’s recent decision to support OpenAI data center purchases was seen as a net negative by Wall Street, which is already concerned about AI spending and the sustainability of such transactions.
AMD has also, in the past (and most recently with Anthropic), tied purchases of its own hardware to investments or stakes it would take in artificial intelligence companies, moves that the market welcomed before but which lately might be viewed with a bit more hostility.
The announcement comes at a time when Cerebras might need it more than AMD: Cerebras listed on Nasdaq in May, priced at $185, opened at $350, and closed its first day at $311.07 before falling back to around $227 in late June 2026.
AMD stock, on the other hand, is up 121.48% year-to-date (YTD) as investors continue to bet heavily on its new Instinct AI processors, and the partnership with Cerebras allows it to further consolidate its gains, as this could be seen as another vote of confidence in its current direction from one of its potential customers.
Follow TechRadar on Google News and add us as a preferred source to receive news, reviews and opinions from our experts in your feeds.




