Skip to main content

folkfox

Skip to main content
Skip to content
AI CONSULTANCY

The cheap coding agent that is almost frontier

Cognition published SWE-2 on 10 September with a blunt promise: a coding agent that performs within a point of the frontier at a fraction of the price. The numbers are real. The comparison underneath them is the part that needs reading.

Quick answerA coding agent is a model that plans, edits and verifies code across many steps. Cognition's SWE-2 scores 50.0 on FrontierCode 1.1 Main, within a point of a frontier rival at 64% lower cost, but trails badly on long-horizon work.
SECTION 01

What Cognition actually shipped#

64%

lower cost than a frontier coding rival, on Cognition's own comparison

Cognition, 10 September 2026

On 10 September 2026 Cognition published SWE-2, and the framing was unusually direct. On the company's own figures the model scores 50.0% on FrontierCode 1.1 Main, one point behind Claude Fable 5.1 at 50.9% and 3.3 behind GPT-6 Astra at 53.3%, while costing a claimed 64% less than Fable 5.1 to run. Read Cognition's launch post and the shape of the argument is clear: not the best model, the best value per unit of work.

For anyone buying an ai coding assistant this year, that is the interesting part. The frontier has stopped being a single point that one lab owns, and started being a curve that several labs are walking along at different prices. A quiet quarry is worth more than a loud one, and a cheap capable model changes what a team can afford to attempt.

The published numbers are specific. Cognition reports 92.8% on Terminal-Bench 2.1, which it describes as the highest recorded in that published set, and 27.3% on Terminal-Bench 4.0, which is where the story turns. Fable 5.1 scores 55.8% and GPT-6 Astra 57.9% on the same long-horizon suite, so the cheap model loses roughly half its capability exactly where the tasks get long. The BenchLM tracker carries the same headline score across the suites it monitors.

@omarsar0
Just watch that massive cost reduction. SWE-2 achieves 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1, while being 64% cheaper.
10 September 2026View on X

The mechanism is worth one paragraph, because it explains why this keeps happening. SWE-2 was post-trained on Kimi K3, a 2.8 trillion parameter mixture-of-experts base, and Cognition scaled reinforcement learning across trillions of parameters rather than building a larger model from scratch. The open-weight release behind that base is why a smaller lab can walk the same trail as a frontier one. Effort levels, low through high, let a caller decide how much reasoning to buy per task, and the clash of pricing that developers were tracking the same day is the undergrowth this dial grows in. That dial is the commercial story, and it is the thing a growth team can actually use.

SECTION 02

The vendor owns the benchmark#

Every lab publishes tables in which its own model wins, and Cognition is no exception: FrontierCode is Cognition's own suite, and the 64% saving is a ratio the company calculated. None of that makes the numbers false. It does mean every figure in this article has to sit next to the name of whoever published it, which is why they do.

Independent trackers agree on the headline. Public ai benchmarks and vendor suites rarely disagree as sharply as the argument around them suggests, and here they line up: OfficeChai reported the 50.0% score and the cheaper-than-frontier framing the same day, and GoKawiil walked the serving stack and the terminal-bench spread, including the 4.0 weakness. The FrontierSWE evaluation repository publishes the methodology behind the suite. Where the independent reading differs from the launch post is not the score but the emphasis: the gap on long-horizon work gets a sentence in the announcement and a paragraph in the analysis.

Capability against relative cost
Scatter chart plotting pass rate against relative cost per task for seven coding agent models, showing the SWE-2 coding agent close to the frontier at the lowest cost1.510.50SWE-2 (50, 0.4)SWE-2Fable 5.1 (50.9, 1)Fable 5.1GPT-6 Astra (53.3, 1.4)GPT-6 AstraGPT-5.6 Sol (47.5, 1.1)GPT-5.6 SolGrok 4.6 (48, 0.8)Grok 4.6Kimi K3 (44.2, 0.5)Kimi K3SWE-1.7 (42, 1)SWE-1.7FrontierCode 1.1 Main pass rate (%)
Scatter chart plotting pass rate against relative cost per task for seven coding agent models, showing the SWE-2 coding agent close to the frontier at the lowest cost
ItemValue
SWE-2 (50, 0.4)SWE-2 (50, 0.4)
Fable 5.1 (50.9, 1)Fable 5.1 (50.9, 1)
GPT-6 Astra (53.3, 1.4)GPT-6 Astra (53.3, 1.4)
GPT-5.6 Sol (47.5, 1.1)GPT-5.6 Sol (47.5, 1.1)
Grok 4.6 (48, 0.8)Grok 4.6 (48, 0.8)
Kimi K3 (44.2, 0.5)Kimi K3 (44.2, 0.5)
SWE-1.7 (42, 1)SWE-1.7 (42, 1)
The cheapest capable coding agent sits almost on the frontier for short tasks: SWE-2 within a point of Fable 5.1 at roughly a third of the relative cost, with GPT-6 Astra ahead on score and far behind on price. Relative cost is modelled from vendor-published comparisons, not an invoice.

Read the scatter the way a fox reads a hedgerow rather than a horizon. The frontier is not a wall, it is a ridge, and the cheap model is walking just under it. What the picture cannot show is task length. A point is an average over many tasks, and averages hide the long tail where the promise gets thin.

The long-horizon gap is the honest headline#

Terminal-Bench 4.0 is the suite that runs agents for hours rather than minutes, and it is the one where SWE-2 loses half its capability: 27.3% against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. That matters because agentic coding in production is mostly long-horizon work. A model that plans beautifully for twelve steps and then forgets its own test suite has not saved you anything, it has moved the cost from the invoice to the reviewer's calendar.

SECTION 03

What a coding agent really costs you#

Headline price per million tokens is not a bill, and this is where most ai evaluation quietly goes wrong. Caching, context tiering and the number of turns a model needs all move the real number, and Cognition's own comparison leans on that: at medium effort, SWE-2 finishes evaluations in 58% fewer turns than its predecessor and costs 81% less per task. The context retrieval work published earlier explains the mechanism: an agent that greps its own repository instead of reading all of it spends fewer tokens finding the den.

One generation of efficiency, indexed
One generation of efficiency, indexedSlope chart showing cost per task, turns per task and steps to a first edit each falling from an index of 100 for the older model to a much lower value for the SWE-2 coding agentSWE-1.7SWE-2Cost per task: 100 to 19Cost per task 100%19%Turns per task: 100 to 42Turns per task 100%42%Steps to first edit: 100 to 38Steps to first edit 100%38%
Slope chart showing cost per task, turns per task and steps to a first edit each falling from an index of 100 for the older model to a much lower value for the SWE-2 coding agent
ItemValue
Cost per task100 to 19
Turns per task100 to 42
Steps to first edit100 to 38
Every measure of wasted effort fell together: Cognition reports task cost down 81%, turns down 58% and median steps to a first edit down from 48 to 18, each indexed here to its predecessor at 100.

The median path to a first concrete patch fell from 48 steps to 18, which is the least glamorous number in the launch and probably the most useful. Wasted exploration is what makes an agent feel slow, and it is what makes a finance director ask why the bill moved when nothing shipped. A model that finds the file sooner is cheaper in a way that never appears in a per-token rate.

Three numbers to watch after a switch

Cost per merged change

100%

Index this to your current model before you migrate, or you cannot tell whether the saving is real.

Turns per task

42%

Falls first when a model gets better at orienting itself. Benchmark it weekly.

Long-horizon failure rate

73%

The share of very long tasks the cheap model is reported to miss. Watch this before it watches you.

Two of those three are cheap to instrument. Cost per merged change is a query against your own repository history, and turns per task falls out of the agent logs you are already storing. The third needs a deliberate test set of the jobs that genuinely run for hours, which is the only way to find out whether the 27.3% is somebody else's problem or the reason your release slipped.

SECTION 04

Five checks before a switch#

Model migrations are cheap to start and expensive to unwind, because the surrounding prompts, tools and test harnesses all get tuned to the model you leave behind. Five checks, run in order, will tell you in a fortnight what a vendor table cannot tell you at all.

The five checks, in order
Recreate the eval

Build a fifty-task set from your own merged pull requests, weighted towards the jobs that take hours rather than minutes. Vendor suites measure the vendor.

Price the whole turn

Log tokens, cache reads and turns per task for a week on both models. Compare cost per merged change, never cost per million tokens.

Hunt the failure mode

Run the long tasks until something breaks, then read the transcript. A model that fails loudly is safer than one that fails politely.

Check the data path

Confirm where prompts and repository content are processed, retained and used for training, in writing, before production data goes near it.

Set the exit

Keep the old model routable for thirty days and a one-line switch in config. Cheap models are replaced on a quarterly rhythm.

Each check with the team that owns it and the artefact that shows it was done properly.
CheckOwnerEvidence
Recreate the evalEngineeringA scored task set that runs in CI
Price the whole turnFinance and EngineeringCost per merged change, both models, one week
Hunt the failure modeEngineeringTranscripts of the tasks that failed
Check the data pathLegal and SecurityA written retention and training position
Set the exitEngineeringA config flag and a tested rollback
  • Recreate the evalEngineeringA scored task set that runs in CI
  • Price the whole turnFinance and EngineeringCost per merged change, both models, one week
  • Hunt the failure modeEngineeringTranscripts of the tasks that failed
  • Check the data pathLegal and SecurityA written retention and training position
  • Set the exitEngineeringA config flag and a tested rollback

The first two checks are the ones that change decisions most often, and they are also the two that get skipped. Teams migrate on a benchmark number and then discover that their own workload is unusually long, unusually context-heavy, or unusually dependent on one tool the model was never trained to call.

None of this argues against cheap capable models. It argues against switching on a single number, which is the habit every launch week invites.

SECTION 05

What it means for next year's budget#

The commercial consequence of a launch like this is not that everything gets cheaper. It is that the price of a fixed unit of capability falls while the price of reliability stays where it was. Budgets that were written around one frontier model per team stop making sense when a second model is nearly as good for a third of the cost on the tasks you run most.

A practical budget line for 2027 therefore has three parts rather than one. A baseline model for everyday work, a premium model reserved for the tasks that run for hours, and a small recurring evaluation cost that keeps both honest. Teams that skip the third part end up re-litigating the same migration every quarter on the strength of whoever published the most recent table.

The frontier is not a place one lab owns. It is a price, and the price is moving.
folkfox, on buying ai evaluation rather than benchmarks

This is the work folkfox does under AI consultancy: model selection on your workload rather than a vendor's, an evaluation harness that survives the next launch, and a cost model your finance team can argue with. It sits alongside the SEO and GEO work, because the same visibility logic governs which answers reach your buyer.

The vulpine reading of a week like this one is straightforward. Cognition shipped a genuine improvement, the independents broadly agree, and the honest caveat is printed in the same table as the win. A buyer who reads all three of those lines will make a better decision than one who reads only the first, and the trail from a benchmark to a real budget is exactly where the next advantage is waiting.

Questions

Frequently asked questions#

What is a coding agent?

A coding agent is a model that works in steps rather than single answers: it reads a repository, plans a change, edits files, runs tests and iterates. That loop is why cost per task matters more than the price per million tokens, since a model that needs fewer turns can be far cheaper in practice.

How much does a coding agent cost per task?

It depends on turns, context and caching rather than the headline rate. Cognition reports SWE-2 at 81% lower cost per task than its previous model at medium effort, driven mostly by finishing in 58% fewer turns. The only reliable figure is one you measure on your own workload.

Are vendor benchmarks trustworthy?

They are usually accurate about the number and selective about the emphasis. Cognition's 50.0% on its own FrontierCode suite is reported consistently by independent trackers, while the long-horizon weakness on Terminal-Bench 4.0 gets far less space. Attribute every figure to whoever published it and read the second table as closely as the first.

Should we switch models to save money?

Not on a launch-week number. Build a fifty-task evaluation from your own merged changes, measure cost per merged change on both models for a week, and keep the old model routable for thirty days. Most teams that do this end up routing rather than switching.

How do you evaluate an ai coding assistant?

Score it on your work, not on a public suite: cost per merged change, turns per task, and the failure rate on the longest jobs you run. Add a written position on where prompts and repository content are processed, and a tested rollback path before anything reaches production.

What is the difference between SWE-2 and a frontier model?

Roughly one point of accuracy on short coding tasks and a much larger difference on long ones. SWE-2 scores 50.0% against Fable 5.1's 50.9% on FrontierCode 1.1 Main at 64% lower cost, but 27.3% against 55.8% on Terminal-Bench 4.0, which is the suite that runs agents for hours.

Keep reading

Read more on this topic#

Buying a cheaper model on someone else's benchmark?

folkfox builds evaluation harnesses and cost models that measure a coding agent on your workload, so a launch-week claim becomes a decision you can defend.

Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.