Nvidia paid $20 billion for SRAM decoding; instead, AMD partnered for it



  • AMD and Cerebras will split inference across two machines, with Helios racks handling fast processing and the Wafer-Scale engine generating tokens, available through Cerebras Cloud in the second half of 2026.
  • Nvidia is also doing something similar by licensing SRAM decoding technology from Groq, an AI chip startup, for $20 billion.
  • The move sees AMD and Cerebras claim 5x higher tokens per watt compared to a standalone Cerebras WSE setup.

AMD and Cerebras Systems have announced a technical partnership that combines the former’s Helios rackscale system with the latter’s Wafer-Scale Engine in what both companies call a disaggregated inference solution.

The move has enabled a combined AMD Helios and Cerebras WSE configuration to deliver up to five times the tokens per second per watt (TPS/W) in internal tests conducted by both chip designers.

Leave a Comment

Your email address will not be published. Required fields are marked *