FrontierCode Benchmark Launches with Claude Opus 4.8 Leading Performance
- Published
- Sep 17, 2026 — 14:49 UTC
FrontierCode Benchmark Overview
FrontierCode has been introduced as a new benchmark for measuring code quality and mergeability, consisting of 150 tasks.
Claude Opus 4.8 achieved a score of 34.3% on the FrontierCode Main subset and 51.8% on the Extended subset, making it the best-performing model. In contrast, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%. The best-performing open-source model, Kimi K2.6, scored 16% on the Main subset and 3.8% on the Diamond subset.
The benchmark features a Diamond subset of 50 hardest tasks and a Main subset of 100 hardest tasks. Developers reportedly spend an average of 40 hours on each task, with FrontierCode achieving an 81% lower false positive rate compared to the previous SWE-Bench Pro benchmark.
Cognition, the organization behind the evaluation process, has collaborated with industry leaders such as Tomer Nosrati, CEO of Celery; Martin McKeaveney, Co-Founder of Budibase; Merlijn Vos, Core Maintainer of uppy; and Claudio Costa, Core Maintainer of Mattermost.
Tomer Nosrati noted the privilege of working with the FrontierCode team, while Martin McKeaveney emphasized the benchmark's attention to detail. Merlijn Vos stated that FrontierCode represents a milestone for AI models in terms of subjective quality evaluation, and Claudio Costa highlighted the meticulous care in the benchmark's criteria.
This follows METR's claims from March 10, 2026, indicating that many SWE-bench-passing pull requests would not be merged into the main branch, underscoring the relevance of FrontierCode in addressing quality in software engineering evaluations.
By Turing Wire Newsdesk · Sep 17, 2026 · How we work →
Summarised from Cognition Labs Blog's original report by the Turing Wire Newsdesk. Read the original for the full story.
Source: Cognition Labs Blog
