Notable other NVIDIA

Nvidia's Groq 3 LPX Claims 4x Speed Over Cerebras with 64 Accelerators

Published
Aug 25, 2026 — 12:27 UTC
Also in this story: Cerebras Google DeepMind

Nvidia’s Groq 3 LPX inference chip claims a performance of 3,400 tokens per second on the Gemma 4 31B model, asserting it is four times faster than Cerebras. However, achieving this performance necessitates a minimum of 64 accelerators, while Cerebras can reach comparable performance with only one or two accelerators. This discrepancy highlights the complexity behind Nvidia’s performance claims. Groq 3 LPX is currently moving into full production, although no specific timeline has been provided. This follows Nvidia’s previous announcements regarding its AI hardware capabilities, including the recent rise in AI server prices due to a DRAM shortage. Practitioners should consider the hardware requirements when evaluating the Groq 3 LPX’s performance against competitors like Cerebras, as the total cost of deployment may significantly differ. For more details, see The Decoder.

Turing Wire

By Callan Zhang · Aug 25, 2026 · Editorial standards →

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: The Decoder