Nvidia's Groq 3 LPX Claims 4x Speed Over Cerebras with 64 Accelerators
- Published
- Aug 25, 2026 — 12:27 UTC
Nvidia’s Groq 3 LPX inference chip claims a performance of 3,400 tokens per second on the Gemma 4 31B model, asserting it is four times faster than Cerebras. However, achieving this performance necessitates a minimum of 64 accelerators, while Cerebras can reach comparable performance with only one or two accelerators. This discrepancy highlights the complexity behind Nvidia’s performance claims. Groq 3 LPX is currently moving into full production, although no specific timeline has been provided. This follows Nvidia’s previous announcements regarding its AI hardware capabilities, including the recent rise in AI server prices due to a DRAM shortage. Practitioners should consider the hardware requirements when evaluating the Groq 3 LPX’s performance against competitors like Cerebras, as the total cost of deployment may significantly differ. For more details, see The Decoder.
By Callan Zhang · Aug 25, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: The Decoder