AMD Cerebras inference platform targets Nvidia’s Groq LPU approach
AMD will pair Instinct GPUs with Cerebras wafer-scale accelerators for low-latency AI inference, with cloud availability planned later this year.
By Dominic Okoye · Staff Writer
· 3 min read
AMD and Cerebras Systems announced an AMD Cerebras inference collaboration that will combine AMD Instinct GPUs with Cerebras’ SRAM-based AI accelerators in a disaggregated compute platform. The companies said the system is aimed at low-latency inference for agentic AI workloads, a part of the AI stack where memory speed can matter as much as raw compute.
The partnership was announced Thursday during AMD CEO Lisa Su’s Advancing AI keynote. Financial terms were not disclosed, and neither company provided specific benchmark results for the combined system. The companies did say they expect the architecture to improve tokens generated per second per watt by as much as 5x.
The design separates the work normally handled inside a single accelerator cluster. AMD’s Instinct GPUs would run compute-heavy prompt processing, while Cerebras’ wafer-scale engines would handle the memory-intensive token generation step. The companies said that split is intended to raise interactivity while preserving throughput and cost efficiency.
What is AMD and Cerebras building?
AMD and Cerebras are building a disaggregated inference platform that uses different accelerators for different parts of model serving. In practical terms, AMD supplies the GPU and rack infrastructure, including Instinct and Helios, while Cerebras supplies wafer-scale engines that use on-chip SRAM rather than high-bandwidth memory.
Cerebras’ pitch is that SRAM on the accelerator can deliver much faster memory access than external memory approaches used by GPUs. Its wafer-scale engines do not depend on HBM4, and Cerebras’ inference systems have been described as reaching output speeds above 2,000 tokens per second in some deployments.
Cerebras CEO and co-founder Andrew Feldman framed the combination as a way to pair AMD’s memory capacity with Cerebras’ memory bandwidth. “What you have with Instinct and the Helios rack is you have the leader in performance and memory capacity,” Feldman said on stage. He said adding Cerebras’ Wafer Scale Engine creates a system he called “unmatched.”
How does this compare with Nvidia and Groq?
The move puts AMD and Cerebras closer to Nvidia’s recent inference strategy around Groq-style language processing units. Cerebras’ accelerators are intended to serve a similar role to Groq 3 LPUs, which were announced alongside Nvidia’s Vera Rubin rack systems at GTC in March.
The companies’ positioning is that Cerebras can serve very large models with far fewer SRAM-based accelerators than an LPU-heavy design. For a trillion-parameter model such as Kimi K2.5, the comparison given was roughly two thousand Groq LPUs for Nvidia’s approach versus at most a few dozen Cerebras accelerators for the AMD-Cerebras system.
That comparison will need public benchmarks before buyers can judge it. The companies did not disclose pricing, deployment configurations, customer commitments, model coverage, or measured latency numbers for the joint platform.
The combined offering is expected to become available through Cerebras Cloud later this year. Su also suggested the Cerebras work may not be a one-off. During a press conference after the keynote, she said AMD’s open ecosystem approach means it will work with multiple companies that have useful workload-specific acceleration technology. “You can expect that we're going to do more workload disaggregation going forward,” Su said.
This story draws on original reporting from The Register.