Models & agents

Cognition ships SWE-2 in Devin, saying it scores within a point of Fable 5.1 on FrontierCode and costs 64 percent less

September 10, 2026 at 10:20 AM PT

Cover image for Cognition's SWE-2 announcement blog post

Image: Cognition

Why it mattersA frontier-adjacent coding model priced well below the leader shifts the cheap-and-good tier, so teams already paying for SWE-1.7 have a decision to make on their next billing cycle.

Cognition released SWE-2 today, the follow-up to its SWE-1.7 coding model, and put it in Devin Desktop and the CLI with a rollout to Devin Web and Fusion in progress. Cognition says SWE-2 was post-trained on Kimi K3, which it describes as a 2.8-trillion-parameter base model that had already been reinforcement-learned for agentic coding.

The numbers Cognition is claiming

On FrontierCode 1.1 Main, Cognition says SWE-2 scores 50.0 percent, against 50.9 percent for Fable 5.1 and 53.3 percent for GPT-6 Astra, with SWE-1.7 at 42.0 percent on the same benchmark. On DeepSWE 1.1 the company reports 73.0 percent for SWE-2 and 74.1 percent for GPT-6 Astra, with SWE-1.7 at 37.7 percent. On Terminal-Bench 2.1 Cognition puts SWE-2 at 92.8 percent, up from 81.5 percent for SWE-1.7. Every one of these is a vendor benchmark from Cognition's own release post, so treat the comparisons as marketing until an independent harness replicates them.

The cost claim is where the news is. Cognition says SWE-2 scores "within one point" of Fable 5.1 on FrontierCode while costing "64 percent" less to run, and that SWE-2 medium beats SWE-1.7 on the same benchmark while taking "58 percent fewer turns" and costing "81 percent less on average". The model exposes three effort levels: medium is described as quicker action for simple and intermediate tasks, and high and max as more planning and exploration for complex ones.

What actually changes for a Devin user

If your team already runs Devin on SWE-1.7 for day-to-day coding, SWE-2 is a drop-in that Cognition claims halves the turn count and drops the bill by around four fifths on the same jobs. If you are paying a Fable 5.1 or an Astra price for the ceiling and only occasionally need it, medium and high on SWE-2 are pitched as most of the score at a smaller fraction of the cost, with max available for the harder work. Cognition is not shipping public pricing in the post, so the concrete number will come from your account, not from the blog.

The Fusion product, which routes work between models in a Devin session, is still rolling out with SWE-2 as one of the options, and Devin Web is being updated. Devin Desktop and the CLI have it available immediately.

The judgement worth holding until an outside test lands is the Fable 5.1 comparison. A model priced 64 percent below a frontier competitor while scoring within a point of it on the vendor's own benchmark is exactly the kind of claim that a real workload rearranges within a week, and Terminal-Bench 4 is where SWE-2 falls to 27.3 percent in Cognition's own table, so the ceiling behaviour is not the same as the frontier's yet. Read the release post for the caveats.

Source

SWE-2 on the Cognition blog, 10 September 2026. Hacker News discussion.

Source: Cognition

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

GameForgeBench evaluates coding agents on building and repairing games in Godot, Unity, Roblox, Minecraft, Unreal and the web

GamePhanes Studio has published GameForgeBench, a Harbor-compatible benchmark that judges coding agents on the full loop of importing, building, running, probing and repairing interactive games in real game engines.

Source: GitHubDev tools

Thesys releases OUI-1, a 4B-active DiffusionGemma finetune that writes UI code and runs on a consumer GPU

Thesys, the team behind OpenUI Lang, released OUI-1, a DiffusionGemma finetune with 26B parameters and 4B active per token, trained to write user-interface code that the OpenUI parser accepts on the first try.

Source: Hacker NewsModels & agents

Kuber Mehta argues Minecraft-in-one-prompt and the pelican-on-a-bicycle SVG are demo benchmarks that labs plainly optimise for by the next launch

A short essay by Kuber Mehta, at 73 Hacker News points on 6 September, argues that the viral one-prompt tests that follow every model launch are fixed public targets a lab has eight weeks to overfit, and picks holdout evals as the alternative that still holds up.

Source: Hacker NewsModels & agents