NotableotherNVIDIA

Nvidia's Groq 3 LPX Claims 4x Speed Over Cerebras with 64 Accelerators

Published
Aug 25, 2026 12:27 UTC
Also in this story:CerebrasGoogle DeepMind

Nvidia's Groq 3 LPX inference chip claims a performance of 3,400 tokens per second on the Gemma 4 31B model, asserting it is four times faster than Cerebras. However, achieving this performance necessitates a minimum of 64 accelerators, while Cerebras can reach comparable performance with only one or two accelerators. This discrepancy highlights the complexity behind Nvidia's performance claims. Groq 3 LPX is currently moving into full production, although no specific timeline has been provided. This follows Nvidia's previous announcements regarding its AI hardware capabilities, including the recent rise in AI server prices due to a DRAM shortage. Practitioners should consider the hardware requirements when evaluating the Groq 3 LPX's performance against competitors like Cerebras, as the total cost of deployment may significantly differ. For more details, see The Decoder.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: The Decoder