AMD Fires Back At NVIDIA’s Groq Bet, Fuses The Cerebras Wafer-Scale Engine With Helios For 5x Higher Tokens Per Second Per Watt

Rohail Saleem
Logos of AMD and Cerebras are displayed on a black background.

NVIDIA scooped up Groq as soon as its LPU showed promise in terms of efficient inferencing. Now, AMD has countered NVIDIA's gambit by partnering with Cerebras to integrate its Helios rack-scale solution with Cerebras' Wafer-Scale Engine, dramatically increasing the inferencing capabilities of the integrated system.

LPU vs. Wafer-Scale Engine

For the benefit of those who might not be aware, Groq's Language Processing Unit (LPU) clusters hundreds or even thousands of specialized chips together, where each individual chip contains giant blocks of Matrix Multiply (MXM) and Vector (VXM) units as well as around 230MB of blazing-fast SRAM. Also, AI model weights are hard-baked directly into the SRAM, completely bypassing the concept of a memory cache.

Related Story AMD Says It Now Controls Nearly Half Of The Data Center CPU Market, And Its Total Compute TAM Will Reach $2 Trillion By 2030

Crucially, the LPU has no branch predictors or hardware schedulers. Instead, the Groq compiler plans every single calculation down to the exact nanosecond, ensuring that relevant data arrives from the SRAM for processing in a continuous, meticulously planned operational cadence, resulting in extremely fast inferencing.

In contrast, Cerebras' Wafer-Scale Engine places an entire AI supercomputer’s worth of memory and compute onto a single, giant, interconnected sheet of silicon, where hundreds of thousands of compute cores and tens of GBs of SRAM connect seamlessly, allowing data to move efficiently without ever hitting external network bottlenecks.

The LPU uses small isolated pools of the SRAM, where a given AI model is broken apart and then distributed across a large number of specialized chips. Cerebras' Wafer-Scale Engine, however, can hold an entire medium-sized model - or huge pieces of a large model - within its unified chunk of SRAM. Also, the versatility of Cerebras' approach enables training as well as inferencing workloads.

AMD is integrating its rack-scale Helios offering with Cerebras' Wafer-Scale Engine to deliver "the ultra-low latency required for the most advanced AI applications"

Under AMD's envisioned roadmap, Helios will provide a high-performance, scalable throughput engine, while Cerebras' Wafer-Scale Engine technology will provide ultra-fast, ultra-low latency decode and token generation. Together, the two compute engines are expected to deliver up to 5x higher tokens per second per watt.

This arrangement balances the high token generation requirements of volume-based workloads with faster response times prized by coding and other agentic tasks.

AMD goes on to note:

"Helios provides ultra-high throughput, processing prompts and large context windows. The Cerebras Wafer-Scale Engine accelerates the memory-bandwidth-intensive token generation, with ultra-low latency. By connecting these best-in-class engines through one integrated workflow, the companies are creating a differentiated platform for ultra-low-latency inference without sacrificing throughput or scale."

Of course, NVIDIA is now selling its LPU-based rack-scale offering, dubbed the Groq 3 LPX, featuring 256 interconnected Groq 3 LPU accelerators, full liquid cooling, and 315 PFLOPS of inference power.

Even so, the AMD-Cerebras collaboration offers a rack-scale product that is much more versatile than the relatively rigid LPUs within the Groq 3 LPX.

Rohail Saleem Photo

About the author: Writing is my one incontrovertible passion. Over the past six years, he has authored over 2,200 distinct articles on financial and tech-related topics, spanning nearly 1 million words. And he has been a member of Wcctech mobile team since 2025. As an alumnus of the University of Toronto, Rotman Commerce Program, I bring nuance, in-depth knowledge, and a unique perspective to every topic that I cover. When I'm not writing, I'm traveling the world, exploring hidden confectionaries and restaurants as an aspiring food connoisseur.

Follow Wccftech on Google to get more of our news coverage in your feeds.

Button