Anthropic's Claude Sonnet 5.5 Matches Opus 5.5 on Benchmarks, Costs 30% Less
- Published
- Sep 28, 2026 — 18:02 UTC
Claude Sonnet 5.5, released on September 28, 2026, achieves a GDPval-AA score of 1,844, nearly matching Opus 5.5's score of 1,846. Anthropic claims Sonnet 5.5 generates output more than 30% faster than its predecessor, Sonnet 5, which scored only 1,449 on the same benchmark. Additionally, Sonnet 5.5 is reported to cut per-task costs by up to 30%, with input token costs set at $2 per million tokens and output token costs at $10 per million tokens.
In benchmark tests, Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, a significant improvement from Sonnet 5's 10.3%. Similarly, it scored 55.5% on CursorBench 4.0 compared to Sonnet 5's 34.1%. The FrontierCode 1.1 benchmark showed Sonnet 5.5 at 46.2% versus 42.4% for Sonnet 5, and it achieved a 61.6% score on Chartography, up from 15.6% for Sonnet 5.
Sonnet 5.5 is also noted for its ability to play through Pokémon Red using only screenshots, marking a significant advancement in interactive AI capabilities. Requests involving high-risk cybersecurity tasks are still routed to Sonnet 5, indicating ongoing optimization challenges. The upcoming Claude Haiku 5.5 is expected to build on these advancements in the coming weeks.
By Turing Wire Newsdesk · Sep 28, 2026 · How we work →
Summarised from The Decoder's original report by the Turing Wire Newsdesk. Read the original for the full story.
Source: The Decoder
