Nvidia's SoL-Pi System Reduces Token Usage by Up to 54.3% Compared to Competitors
- Published
- Sep 26, 2026 — 10:30 UTC
Nvidia's SoL-Pi System Reduces Token Usage by Up to 54.3% Compared to Competitors
Nvidia's SoL-Pi system achieves a 50% reduction in token usage compared to Codex and a 54.3% reduction compared to Claude Code. This system, which costs $894, is a significant improvement over the $1,339 Pi system, offering estimated savings of $8.75 to $13.50 per hour when using Codex and Claude Code, and $4.36 to $5.71 per hour with Pi.
SoL-Pi operates across 535 executable environments, generating 3,000 total runs and facilitating 60,000 agent-environment interactions. In tests on the Terminal-Bench 4, SoL-Pi solved 15 tasks, while Codex and Pi each solved 18 tasks. Performance retention on the Opus 5 model is reported at 94.3% of Pi's performance.
The system's efficiency gains are attributed to recursive self-improvement, although Nvidia researchers caution that shorter context lengths can reduce prompt cache reuse. The average user instructions preserved by compression is 17%, and nearly 70% of token usage comes from cached prompts.
Peter Walker, an analyst, noted that the cost per solved task varied by nearly 3x, despite the same model performing the tasks. Eric Provencher, a developer, emphasized that using more than two sub-agents typically increases token consumption without improving output quality. This follows a trend of 14x growth in agentic token usage since February 2026, highlighting the need for optimization in AI coding agents.
By Callan Zhang · Sep 26, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: The Decoder
