

The cheap coding agent that is almost frontier
Cognition published SWE-2 on 10 September with a blunt promise: a coding agent that performs within a point of the frontier at a fraction of the price. The numbers are real. The comparison underneath them is the part that needs reading.
By Katie Delaney / 2026-09-10 / 12 min read

What Cognition actually shipped#
lower cost than a frontier coding rival, on Cognition's own comparison
On 10 September 2026 Cognition published SWE-2, and the framing was unusually direct. On the company's own figures the model scores 50.0% on FrontierCode 1.1 Main, one point behind Claude Fable 5.1 at 50.9% and 3.3 behind GPT-6 Astra at 53.3%, while costing a claimed 64% less than Fable 5.1 to run. Read Cognition's launch post and the shape of the argument is clear: not the best model, the best value per unit of work.
For anyone buying an ai coding assistant this year, that is the interesting part. The frontier has stopped being a single point that one lab owns, and started being a curve that several labs are walking along at different prices. A quiet quarry is worth more than a loud one, and a cheap capable model changes what a team can afford to attempt.
The published numbers are specific. Cognition reports 92.8% on Terminal-Bench 2.1, which it describes as the highest recorded in that published set, and 27.3% on Terminal-Bench 4.0, which is where the story turns. Fable 5.1 scores 55.8% and GPT-6 Astra 57.9% on the same long-horizon suite, so the cheap model loses roughly half its capability exactly where the tasks get long. The BenchLM tracker carries the same headline score across the suites it monitors.
Just watch that massive cost reduction. SWE-2 achieves 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1, while being 64% cheaper.
The mechanism is worth one paragraph, because it explains why this keeps happening. SWE-2 was post-trained on Kimi K3, a 2.8 trillion parameter mixture-of-experts base, and Cognition scaled reinforcement learning across trillions of parameters rather than building a larger model from scratch. The open-weight release behind that base is why a smaller lab can walk the same trail as a frontier one. Effort levels, low through high, let a caller decide how much reasoning to buy per task, and the clash of pricing that developers were tracking the same day is the undergrowth this dial grows in. That dial is the commercial story, and it is the thing a growth team can actually use.
The vendor owns the benchmark#
Every lab publishes tables in which its own model wins, and Cognition is no exception: FrontierCode is Cognition's own suite, and the 64% saving is a ratio the company calculated. None of that makes the numbers false. It does mean every figure in this article has to sit next to the name of whoever published it, which is why they do.
Independent trackers agree on the headline. Public ai benchmarks and vendor suites rarely disagree as sharply as the argument around them suggests, and here they line up: OfficeChai reported the 50.0% score and the cheaper-than-frontier framing the same day, and GoKawiil walked the serving stack and the terminal-bench spread, including the 4.0 weakness. The FrontierSWE evaluation repository publishes the methodology behind the suite. Where the independent reading differs from the launch post is not the score but the emphasis: the gap on long-horizon work gets a sentence in the announcement and a paragraph in the analysis.
| Item | Value |
|---|---|
| SWE-2 (50, 0.4) | SWE-2 (50, 0.4) |
| Fable 5.1 (50.9, 1) | Fable 5.1 (50.9, 1) |
| GPT-6 Astra (53.3, 1.4) | GPT-6 Astra (53.3, 1.4) |
| GPT-5.6 Sol (47.5, 1.1) | GPT-5.6 Sol (47.5, 1.1) |
| Grok 4.6 (48, 0.8) | Grok 4.6 (48, 0.8) |
| Kimi K3 (44.2, 0.5) | Kimi K3 (44.2, 0.5) |
| SWE-1.7 (42, 1) | SWE-1.7 (42, 1) |
Read the scatter the way a fox reads a hedgerow rather than a horizon. The frontier is not a wall, it is a ridge, and the cheap model is walking just under it. What the picture cannot show is task length. A point is an average over many tasks, and averages hide the long tail where the promise gets thin.
The long-horizon gap is the honest headline#
Terminal-Bench 4.0 is the suite that runs agents for hours rather than minutes, and it is the one where SWE-2 loses half its capability: 27.3% against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. That matters because agentic coding in production is mostly long-horizon work. A model that plans beautifully for twelve steps and then forgets its own test suite has not saved you anything, it has moved the cost from the invoice to the reviewer's calendar.

What a coding agent really costs you#
Headline price per million tokens is not a bill, and this is where most ai evaluation quietly goes wrong. Caching, context tiering and the number of turns a model needs all move the real number, and Cognition's own comparison leans on that: at medium effort, SWE-2 finishes evaluations in 58% fewer turns than its predecessor and costs 81% less per task. The context retrieval work published earlier explains the mechanism: an agent that greps its own repository instead of reading all of it spends fewer tokens finding the den.
| Item | Value |
|---|---|
| Cost per task | 100 to 19 |
| Turns per task | 100 to 42 |
| Steps to first edit | 100 to 38 |
The median path to a first concrete patch fell from 48 steps to 18, which is the least glamorous number in the launch and probably the most useful. Wasted exploration is what makes an agent feel slow, and it is what makes a finance director ask why the bill moved when nothing shipped. A model that finds the file sooner is cheaper in a way that never appears in a per-token rate.
Cost per merged change
Index this to your current model before you migrate, or you cannot tell whether the saving is real.
Turns per task
Falls first when a model gets better at orienting itself. Benchmark it weekly.
Long-horizon failure rate
The share of very long tasks the cheap model is reported to miss. Watch this before it watches you.
Two of those three are cheap to instrument. Cost per merged change is a query against your own repository history, and turns per task falls out of the agent logs you are already storing. The third needs a deliberate test set of the jobs that genuinely run for hours, which is the only way to find out whether the 27.3% is somebody else's problem or the reason your release slipped.
Five checks before a switch#
Model migrations are cheap to start and expensive to unwind, because the surrounding prompts, tools and test harnesses all get tuned to the model you leave behind. Five checks, run in order, will tell you in a fortnight what a vendor table cannot tell you at all.
Build a fifty-task set from your own merged pull requests, weighted towards the jobs that take hours rather than minutes. Vendor suites measure the vendor.
Log tokens, cache reads and turns per task for a week on both models. Compare cost per merged change, never cost per million tokens.
Run the long tasks until something breaks, then read the transcript. A model that fails loudly is safer than one that fails politely.
Confirm where prompts and repository content are processed, retained and used for training, in writing, before production data goes near it.
Keep the old model routable for thirty days and a one-line switch in config. Cheap models are replaced on a quarterly rhythm.
| Check | Owner | Evidence |
|---|---|---|
| Recreate the eval | Engineering | A scored task set that runs in CI |
| Price the whole turn | Finance and Engineering | Cost per merged change, both models, one week |
| Hunt the failure mode | Engineering | Transcripts of the tasks that failed |
| Check the data path | Legal and Security | A written retention and training position |
| Set the exit | Engineering | A config flag and a tested rollback |
- Recreate the evalEngineeringA scored task set that runs in CI
- Price the whole turnFinance and EngineeringCost per merged change, both models, one week
- Hunt the failure modeEngineeringTranscripts of the tasks that failed
- Check the data pathLegal and SecurityA written retention and training position
- Set the exitEngineeringA config flag and a tested rollback
The first two checks are the ones that change decisions most often, and they are also the two that get skipped. Teams migrate on a benchmark number and then discover that their own workload is unusually long, unusually context-heavy, or unusually dependent on one tool the model was never trained to call.
None of this argues against cheap capable models. It argues against switching on a single number, which is the habit every launch week invites.

What it means for next year's budget#
The commercial consequence of a launch like this is not that everything gets cheaper. It is that the price of a fixed unit of capability falls while the price of reliability stays where it was. Budgets that were written around one frontier model per team stop making sense when a second model is nearly as good for a third of the cost on the tasks you run most.
A practical budget line for 2027 therefore has three parts rather than one. A baseline model for everyday work, a premium model reserved for the tasks that run for hours, and a small recurring evaluation cost that keeps both honest. Teams that skip the third part end up re-litigating the same migration every quarter on the strength of whoever published the most recent table.
The frontier is not a place one lab owns. It is a price, and the price is moving.
This is the work folkfox does under AI consultancy: model selection on your workload rather than a vendor's, an evaluation harness that survives the next launch, and a cost model your finance team can argue with. It sits alongside the SEO and GEO work, because the same visibility logic governs which answers reach your buyer.
The vulpine reading of a week like this one is straightforward. Cognition shipped a genuine improvement, the independents broadly agree, and the honest caveat is printed in the same table as the win. A buyer who reads all three of those lines will make a better decision than one who reads only the first, and the trail from a benchmark to a real budget is exactly where the next advantage is waiting.
Frequently asked questions#
What is a coding agent?
A coding agent is a model that works in steps rather than single answers: it reads a repository, plans a change, edits files, runs tests and iterates. That loop is why cost per task matters more than the price per million tokens, since a model that needs fewer turns can be far cheaper in practice.
How much does a coding agent cost per task?
It depends on turns, context and caching rather than the headline rate. Cognition reports SWE-2 at 81% lower cost per task than its previous model at medium effort, driven mostly by finishing in 58% fewer turns. The only reliable figure is one you measure on your own workload.
Are vendor benchmarks trustworthy?
They are usually accurate about the number and selective about the emphasis. Cognition's 50.0% on its own FrontierCode suite is reported consistently by independent trackers, while the long-horizon weakness on Terminal-Bench 4.0 gets far less space. Attribute every figure to whoever published it and read the second table as closely as the first.
Should we switch models to save money?
Not on a launch-week number. Build a fifty-task evaluation from your own merged changes, measure cost per merged change on both models for a week, and keep the old model routable for thirty days. Most teams that do this end up routing rather than switching.
How do you evaluate an ai coding assistant?
Score it on your work, not on a public suite: cost per merged change, turns per task, and the failure rate on the longest jobs you run. Add a written position on where prompts and repository content are processed, and a tested rollback path before anything reaches production.
What is the difference between SWE-2 and a frontier model?
Roughly one point of accuracy on short coding tasks and a much larger difference on long ones. SWE-2 scores 50.0% against Fable 5.1's 50.9% on FrontierCode 1.1 Main at 64% lower cost, but 27.3% against 55.8% on Terminal-Bench 4.0, which is the suite that runs agents for hours.
Read more on this topic#
One staffer, 1.5 million tokens, and a meter nobody was watching
The governance side of the same problem: what happens to a budget when an agent runs unattended.
Read the pieceAI consultancyDeepSeek V4 released a model named after its own expiry date
Another cheap fast model, and the small print that decided whether it belonged in production.
Read the pieceAI consultancyMeta Muse can act. The question is what you will allow it to touch
Approval gates and audit trails, the control layer that matters once agents start doing work.
Read the pieceAI consultancyA UK MP wants a superintelligence treaty. The world has four months
Where the procurement rules may be heading while everyone argues about benchmark tables.
Read the pieceBuying a cheaper model on someone else's benchmark?
folkfox builds evaluation harnesses and cost models that measure a coding agent on your workload, so a launch-week claim becomes a decision you can defend.
Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.