Majorinfrastructure computeNVIDIA

CoreWeave and NVIDIA Launch Vera Rubin NVL72 for Agentic AI Workloads

Published
Sep 30, 2026 — 15:00 UTC
Also in this story:CoreWeaveCognition

CoreWeave and NVIDIA Launch Vera Rubin NVL72 for Agentic AI Workloads

CoreWeave has received its first production racks of the NVIDIA Vera Rubin NVL72, which features 128 CPUs and a total of 11,264 cores per rack. This infrastructure is designed to support agentic AI workloads, providing a 4.8x increase in total token throughput for SWE-2 inference workloads compared to the previous GB200 NVL72.

Cognition, an applied AI lab and the first customer of the Vera Rubin NVL72, has scaled to thousands of GPUs on CoreWeave in just nine months. Devin AI, a software engineer at Cognition, emphasized the importance of agentic coding, which involves complex workloads characterized by long contexts and high concurrency. Silas Alberti from Cognition noted that cost per token is a critical factor in shipping products.

The integration of CoreWeave's Forge environment allows for training and improving AI models, while the CoreWeave Kubernetes Service facilitates the management of containerized applications. CoreWeave ARIA, now generally available, provides tools for analyzing AI runs and proposing experiments, and CoreWeave Sandboxes offer isolated environments for running AI tasks.

NVIDIA's Vera CPU demonstrates a 3x faster agent sandbox startup time and a 1.7x performance gain on Terminal-Bench across all passing tasks. Additionally, serverless reinforcement learning setups using Vera CPUs are 1.4x faster and 40% lower in cost compared to self-managed alternatives.

CoreWeave's ranking as Platinum in SemiAnalysis ClusterMAX 1.0, 2.0, and 3.0 reflects its strong performance in MLPerf results across training and inference. The ongoing Fully Connected event in San Francisco showcases these advancements, emphasizing the collaboration between NVIDIA and CoreWeave engineers to tackle complex problems in AI infrastructure.

Summarised from NVIDIA Blog's original report by the Turing Wire Newsdesk. Read the original for the full story.

Source: NVIDIA Blog