Majormodel releaseCognition

Cognition Labs Launches SWE-2, Achieving 50% on FrontierCode 1.1 Main

Published
Sep 17, 2026 — 14:49 UTC

SWE-2 Launch Overview

Cognition Labs has released SWE-2, achieving a score of 50.0% on the FrontierCode 1.1 Main benchmark. This new model surpasses its predecessor, SWE-1.7, which scored 44.2%, and also outperforms Grok 4.6, which scored 48.0%. SWE-2 is designed to enhance coding efficiency, evidenced by a 64% cost reduction compared to Fable 5.1 and an 81% reduction compared to SWE-1.7 on the same benchmark.

SWE-2 utilizes the Kimi K3 model, which has 2.8 trillion parameters, and is available on platforms including Devin Desktop, Devin Web, and Fusion. The model also shows significant improvements in operational efficiency, taking a median of 18 steps to make its first edit on FrontierCode 1.1 Main, compared to 48 steps for SWE-1.7.

In terms of performance across various benchmarks, SWE-2 scored 73.0% on DeepSWE 1.1 and 92.8% on Terminal-Bench 2.1, while it struggled with Terminal-Bench 4, scoring only 27.3%. The model also passed 98.0% of attempts overall, indicating high reliability.

This follows the July 2026 publication of FrontierCode 1.1 and SWE-1.7, marking a significant advancement in coding model capabilities. The Cognition Team claims that SWE-2 is their most advanced coding model yet, demonstrating superior performance and cost efficiency in comparison to previous iterations.

Summarised from Cognition Labs Blog's original report by the Turing Wire Newsdesk. Read the original for the full story.

Source: Cognition Labs Blog