Majorinfrastructure computeNVIDIA

NVIDIA Vera Rubin NVL72 Achieves 3.7x Throughput Over GB300 NVL72 in MLPerf

Published
Sep 16, 2026 15:00 UTC

NVIDIA's Vera Rubin NVL72 achieves up to 3.7 times higher throughput compared to the GB300 NVL72 in the latest MLPerf Inference v6.1 benchmarks. The GB300 NVL72 submission utilized 288 GPUs and demonstrated a 99% scaling efficiency across four racks. In comparison, the Vera Rubin NVL72 showed a 2.5 times higher throughput on the DeepSeek-R1 benchmark and a 30 times improvement in SemiAnalysis AgentX testing. Additionally, it delivered 0.65 720p videos per second on the WAN 2.2 benchmark, with a processing time of 5.7 seconds per video. The Vera Rubin NVL72 also achieved 9 times higher throughput and 7.5 times lower latency than a single node on the same benchmark. NVIDIA noted that software optimizations contributed to a 1.6 times performance increase over the previous MLPerf Inference v6.0 submissions. This performance enhancement indicates that each Vera Rubin NVL72 rack can serve more users and generate greater revenue than a GB300 NVL72 rack. The results were retrieved from MLCommons on September 16, 2026. This follows NVIDIA's ongoing commitment to continuous software development, which enhances performance and features across its platforms.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: NVIDIA Blog