Notableevaluation benchmarksCognition

FrontierCode 1.1 Released with Updated Model Scores and Grading Criteria

Published
Sep 17, 2026 14:49 UTC

FrontierCode 1.1 has been released, featuring updated model scores for Sonnet 5 and Fable 5. This version relaxes 75 overly strict grading criteria out of over 1,000 audited criteria, improving evaluation flexibility. The evaluation report from METR on GPT-5.6 Sol, dated June 26, 2026, indicates that unfair internet use rates have fallen below 1% for the evaluated models. The original FrontierCode was introduced in 2026, and this update follows just one month after the initial release of FrontierCode 1.0. The documentation claims that the relative performances of the models evaluated did not substantially change compared to version 1.0, while adherence to the prompt remains remarkably good. Additionally, the Extended set of FrontierCode now includes a total of 150 tasks, with 50 tasks previously part of the deprecated Diamond set. This update is part of ongoing efforts to address the issue of unfair internet use in software engineering evaluations, as noted in the FrontierCode 1.1 documentation.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: Cognition Labs Blog