Cognition Labs Launches SWE-2, Achieving 50% on FrontierCode 1.1 Main
- Published
- Sep 17, 2026 — 14:49 UTC
SWE-2 Launch Overview
Cognition Labs has released SWE-2, achieving a score of 50.0% on the FrontierCode 1.1 Main benchmark. This new model surpasses its predecessor, SWE-1.7, which scored 44.2%, and also outperforms Grok 4.6, which scored 48.0%. SWE-2 is designed to enhance coding efficiency, evidenced by a 64% cost reduction compared to Fable 5.1 and an 81% reduction compared to SWE-1.7 on the same benchmark.
SWE-2 utilizes the Kimi K3 model, which has 2.8 trillion parameters, and is available on platforms including Devin Desktop, Devin Web, and Fusion. The model also shows significant improvements in operational efficiency, taking a median of 18 steps to make its first edit on FrontierCode 1.1 Main, compared to 48 steps for SWE-1.7.
In terms of performance across various benchmarks, SWE-2 scored 73.0% on DeepSWE 1.1 and 92.8% on Terminal-Bench 2.1, while it struggled with Terminal-Bench 4, scoring only 27.3%. The model also passed 98.0% of attempts overall, indicating high reliability.
This follows the July 2026 publication of FrontierCode 1.1 and SWE-1.7, marking a significant advancement in coding model capabilities. The Cognition Team claims that SWE-2 is their most advanced coding model yet, demonstrating superior performance and cost efficiency in comparison to previous iterations.
By Turing Wire Newsdesk · Sep 17, 2026 · How we work →
Summarised from Cognition Labs Blog's original report by the Turing Wire Newsdesk. Read the original for the full story.
Source: Cognition Labs Blog
