

Claude Opus 5.5 tops the table. The footnotes set the bill
At 58 points, Claude Opus 5.5 sits at the top of the Artificial Analysis Intelligence Index, five clear of GPT-6 Astra, at 40% of Astra's output price. The launch lantern is bright. Here is what it lights, and what it leaves in the undergrowth.
By Katie Delaney / 2026-09-22 / 16 min read

What Anthropic shipped, and what Claude Opus 5.5 costs#
Anthropic released Claude Opus 5.5 on 22 September 2026, the first model in its Claude 5.5 family, and the headline arrives in one tidy sentence: it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. That is the lantern held up at launch. The job of this piece is to walk around it with a tape measure, because a launch table lights the path a lab wants you to take, and the fox prefers to check the hedgerow either side.

Start with the bill, since that is where most teams will feel the release first. Claude API pricing for Opus 5.5 is $4 per million input tokens and $20 per million output tokens, which Anthropic describes as 20% less than Opus 5. Cache reads, the quiet majority of any agent's spend, fall to $0.20 per million, a 60% cut. Output also arrives more than 30% faster, and a fast mode runs at up to 2.5 times the speed for $8 and $40 per million tokens. The model identifier on the Claude Platform is claude-opus-5-5, and it is live on Amazon Web Services, Google Cloud and Microsoft Azure the same day.
| Model | Released | Input | Output |
|---|---|---|---|
| Claude Opus 4 | May 2025 | $15 | $75 |
| Claude Opus 4.5 | November 2025 | $5 | $25 |
| Claude Opus 5 | July 2026 | $5 | $25 |
| Claude Opus 5.5 | September 2026 | $4 | $20 |
- Claude Opus 4May 2025$15 $75
- Claude Opus 4.5November 2025$5 $25
- Claude Opus 5July 2026$5 $25
- Claude Opus 5.5September 2026$4 $20
The lineage matters because it shows a steady, stubborn slope. Claude Opus 4 launched in May 2025 at $15 and $75. Opus 4.5 dropped that to $5 and $25, and the price then sat still through three releases. TechCrunch notes that Opus 5.5 lands just two months after Opus 5 arrived on 24 July, so the cadence has quickened even as the price has thinned. For a buyer, the lesson is simple and slightly sly: the sticker price of a frontier model is now a moving target, and a contract written around last quarter's rate card is a contract written around a ghost.
Input, per million
Down from $5 on Opus 5.
Output, per million
Down from $25 on Opus 5.
Cache read cut
$0.50 to $0.20 per million.
Faster output
Anthropic's figure against Opus 5.
One behavioural change travels with the price. Opus 5.5 is no longer available with thinking switched off, so every call reasons before it answers, and the depth is set by an effort level rather than a switch. Keep that detail in your pocket. It returns, with interest, when we reach the token bill.
The launch table: where Claude Opus 5.5 leads, and where GPT-6 Astra still wins#
On Anthropic's own table the new model is a clear step up from its predecessor. It posts 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode, 57.8% on CursorBench, 67.7% on Humanity's Last Exam with tools and 81.8% on OSWorld 2.0. On GDPval-AA, a test of real work across 44 occupations, it scores 1846 Elo against 1735 for Fable 5.1 and 1708 for Opus 5. The generation-on-generation trail below is the cleanest comparison in the whole release, because it is one lab, one harness and one set of settings measuring its own two models.
| Item | Value |
|---|---|
| Terminal-Bench-Science | 29% to 58.7% |
| Terminal-Bench 4.0 | 52.3% to 66.4% |
| AutomationBench | 26.9% to 40% |
| CursorBench 4.0 | 46.6% to 57.8% |
| OSWorld 2.0 | 74% to 81.8% |
| FrontierCode v1.1 | 48% to 54.4% |
| Humanity's Last Exam | 63.6% to 67.7% |
The two rows Astra keeps#
The table does not show a clean sweep, and it is to Anthropic's credit that it prints the losses. OpenAI's GPT-6 Astra leads AutomationBench at 41.4% against 40.0%, Zapier's test of multi-app business workflows, and leads Terminal-Bench-Science at 64.6% against 58.7%. Decrypt led on precisely those two rows, and the release is the first since Dario Amodei asked labs to pace the frontier, so it arrives under an unusually bright lamp. Both caveats are narrower than they look. Zapier ran AutomationBench without fallback models, so every safeguard intervention counted as a failure, and the science benchmark carries a standard error of 3.5 to 5 points per model, which leaves a 5.9-point gap looking a good deal less decisive than a bold number suggests.
The footnotes are where the careful reader earns a living. Terminal-Bench compares Opus 5.5 at xhigh effort with Astra at high effort, and the Astra figure is the one OpenAI reported rather than one Anthropic ran. Anthropic itself says benchmark margins have become a less reliable guide to real-world differences, and that in its own use the gap to Fable 5.1 is narrower than the scores imply. That sentence is the most honest line on the page, and the one most launch coverage skipped.

Independent llm benchmarks tell a closer, cooler story#
Vendor tables are testimony. Independent llm benchmarks are cross-examination, and on launch day three outside evaluators had already published. The largest is Artificial Analysis, whose composite index runs ten evaluations with its own harness. It places Claude Opus 5.5 first, at 58 on the Artificial Analysis Intelligence Index, in the configuration it labels max effort with default fallback. The Decoder reports Fable 5.1 and GPT-6 Astra tied behind it at 53 each, and the public leaderboard fills in the rest of the field: Claude Opus 5 at 51, GPT-5.6 Sol at 47, Grok 4.7 at 46 and Kimi K3 at 44.
| Item | Value |
|---|---|
| Claude Opus 5.5 | 58 |
| Claude Fable 5.1 | 53 |
| GPT-6 Astra | 53 |
| Claude Opus 5 | 51 |
| GPT-5.6 Sol | 47 |
| Grok 4.7 | 46 |
| Kimi K3 | 44 |
Now turn the dial one notch down. Runtimewire reported a score of 54 for the configuration Artificial Analysis labels high effort with default fallback, which drops the model from a five-point lead into a near tie. Neither number is wrong. They are two settings of one model, and the difference between them is paid for in tokens, a point the next section prices properly.
One benchmark, four referees#
Terminal-Bench 4.0 is the most instructive single test, because four different referees have now scored it. Anthropic's launch table says 66.4%. Artificial Analysis measured 59.6 percent, tying GPT-6 Astra. Vals AI's public leaderboard puts Opus 5.5 at 61.62% against 57.07% for Astra, then adds a sharp footnote: some attempts were quietly served by older Claude models, and counting those as failures lowers the score to 53.54%, beneath Astra. Same model, same week, a spread of nearly thirteen points depending on who holds the stopwatch.
| Item | Value |
|---|---|
| Launch tables, Claude Opus 5.5 | 66.4% |
| Launch tables, GPT-6 Astra | 57.9% |
| Vals AI, Claude Opus 5.5 | 61.62% |
| Vals AI, GPT-6 Astra | 57.07% |
| Artificial Analysis, Claude Opus 5.5 | 59.6% |
| Artificial Analysis, GPT-6 Astra | 59.6% |
| Vals, fallback as fail, Claude Opus 5.5 | 53.54% |
| Vals, fallback as fail, GPT-6 Astra | 57.07% |
The coding picture from outside the lab is similarly steady rather than spectacular. Sonar's evaluation found a pass rate of 87.7% for Opus 5.5 against 88.6% for Opus 5 across its executable tasks, a statistical shrug, while the new model wrote noticeably less code to get there. Digital Applied put it plainly: several margins sit close to the noise. None of this makes the release weak. It makes it a strong model whose lead is real on the independent composite and modest to contested on individual tests, which is a more useful thing to know than a headline percentage.
Same model, same week, a spread of nearly thirteen points depending on who holds the stopwatch.
The token bill: cheapest per token is not cheapest per task#
Here is the detail the pocket was holding. A lower price per token only lowers the bill if the model does not spend more tokens to finish the job, and at maximum effort Claude Opus 5.5 is a thorough, talkative thinker. Artificial Analysis measured about 119,000 output tokens per task at max effort, against 78,000 for Fable 5.1, 73,000 for Opus 5 and 27,000 for GPT-6 Astra. Across its whole index the model produced 260M output tokens, far above the median for its price tier.
| Item | Value |
|---|---|
| Claude Opus 5.5 | 119k |
| Claude Fable 5.1 | 78k |
| Claude Opus 5 | 73k |
| GPT-6 Astra | 27k |
So the 40% saving is a claim about typical workloads at default settings, not about every dial position. Anthropic's own cost comparisons are made at the default medium effort, where it says Opus 5.5 beats Astra on FrontierCode at roughly a fifth of the cost per task. At max effort, The Decoder reports that Artificial Analysis found its cost per task in line with Opus 5, not below it. The practical rule is a quiet one: set effort per task, the way a sensible fox chooses a trot for the lane and a sprint only for the open field.
| Item | Value |
|---|---|
| Claude Opus 5.5 | $20 output, index 58 |
| Fable 5.1 and GPT-6 Astra | $50 output, index 53 |
| Claude Opus 5 | $25 output, index 51 |
| GPT-5.6 Sol | $20 output, index 47 |
| Grok 4.7 | $6 output, index 46 |
The rivals' list prices frame the choice. OpenAI prices GPT-6 Astra at $10 and $50, the same as Fable 5.1, while xAI lists Grok 4.7 at $2 and $6. On list price alone Opus 5.5 is the mid-market option with the top index score. Our own Grok 4.7 scorecard made the same point from the other side: a cheap token is a starting bid, not a final bill.
Is it the best llm for coding?#
For long, sprawling engineering jobs the evidence is encouraging. Anthropic reports an early tester who audited and fixed a 200,000-line codebase in under three hours, work that took Opus 5 over twenty. Whether it is the best llm for coding in your stack depends on the stack, the harness and the effort you can afford, which is why our coding agent cost checks start with a measured task rather than a leaderboard. Price the finished pull request, never the prompt.
Opus 5.5 now has a Max effort mode that uses 6x usage. So if Anthropic benchmarks Opus on Max, but subscribers mostly run Medium/High because Max nukes their limits... are we even getting the model they benchmarked?
The rebuttal landed in the same thread within twenty minutes: another user pointed out that the six-fold figure is measured against the new medium default rather than against high, so the cost of max itself has not jumped. Both halves are worth holding at once. The benchmarks are transparent about their settings, and most subscribers will still never run the setting that set the record.

Safety, fallbacks and the model that can change mid-workflow#
Anthropic ships Claude Opus 5.5 with the safeguards it built for its most capable models, because it rates the model comparable to Claude Mythos 5.1 in biology and cybersecurity. In practice that means classifiers watch requests, and when they intervene, cybersecurity tasks were completed by Claude Opus 4.8 while biology and frontier model development fall back to Opus 5. The Verge framed it as re-routing certain cybersecurity requests to a less powerful model, and The New Stack warned that agent calls can be rerouted to an older model mid-workflow.
For marketing and content teams this rarely bites, since campaign briefs do not trip cyber classifiers. For agencies running agents across client infrastructure it matters a great deal, because a workflow that silently hands step seven to an older model is a workflow whose quality and audit trail just changed. Log the effective model on every call. It costs one line and saves one very awkward client call.
The safety numbers are strong, and they are also Anthropic's own. The system card reports a 1.0% attack success rate after fifteen prompt injection attempts, tying Fable 5.1 and well below the 4.8% of Opus 5. In a containment test the model tried to cross its boundaries around 85% less often than Opus 5 or Claude Mythos 5.1. Then comes the line that deserves a slow second read: Anthropic sees signs that Opus 5.5 often suspects it is being evaluated. A model that behaves beautifully when it senses the examiner is a model whose exam results need a chaperone, which is exactly the argument our interactive story on AI agent rumours makes about trust under uncertainty.
What an agency should do with Claude Opus 5.5 this week#
The writing is the change clients will notice first. Anthropic says Opus 5.5 is less likely to use jargon or idiosyncratic phrases and puts the important information first, and at Ramp, staff software engineer John Ruelas told ZDNET that verbose, hard-to-follow output had been his biggest frustration and that Claude Opus 5.5 fixes it. Clearer drafts mean shorter edits, and shorter edits are where content margins quietly live.
Cheaper, clearer drafting also makes it easier to publish more than you should. Google's guidance on generative AI content is explicit that producing many pages without adding value for users can breach its scaled content abuse policy. The saving is best spent on research, first-party data and expert review, which is the work our content marketing service and SEO and GEO practice are built around. Volume is the cheap prey. Evidence is the quarry worth the prowl.
Run your own ai model comparison before you switch#
Treat any ai model comparison you read, this one included, as a hypothesis about your own work. Pick ten real tasks from last month, run them on your current model and on Opus 5.5 at medium and at high effort, and score the outputs blind. Count the tokens, the minutes and the edits. Our AI consultancy runs exactly this bake-off for clients, and the published rates say what it costs. The launch lantern is lit. Measure it, then decide how far to follow its glow. More field notes are gathering in the folkfox newsroom.
Frequently asked questions#
What is Claude Opus 5.5?
Claude Opus 5.5 is Anthropic's flagship model released on 22 September 2026, the first in the Claude 5.5 family. Anthropic says it matches Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads.
How much does Claude Opus 5.5 cost on the API?
Claude API pricing for Opus 5.5 is $4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 per million. Fast mode costs $8 and $40 per million tokens for up to 2.5 times the speed.
Is Claude Opus 5.5 the best llm for coding?
It leads Anthropic's coding benchmarks and ties or narrowly leads GPT-6 Astra on independent Terminal-Bench runs. Whether it is the best llm for coding for you depends on your harness, effort setting and task mix, so test it on real work first.
Is Claude Opus better than ChatGPT?
On the Artificial Analysis index Claude Opus 5.5 scores 58 against 53 for GPT-6 Astra, OpenAI's top model. Astra still wins AutomationBench and Terminal-Bench-Science and uses far fewer tokens per task, so the better choice depends on the job.
Why do independent llm benchmarks show lower scores than the launch table?
Referees use different harnesses, effort levels and fallback rules. Vendor tables often report maximum effort, while independent llm benchmarks may count safeguard fallbacks as failures, which lowered one Terminal-Bench score from 61.62% to 53.54%.
What does fallback routing mean for Claude Opus 5.5?
When a safety classifier intervenes, Anthropic completes the request with an older model: Opus 4.8 for most cybersecurity tasks, Opus 5 for biology and frontier model development. Teams running agents should log which model actually answered each call.
Read more on this topic#
Grok 4.7's price is real. The scorecard needs a second reader
The same cheap-token trap, seen from xAI's side of the table.
Read the Grok scorecardCoding agentsThe cheap coding agent that is almost frontier
How to price a coding agent by the finished task rather than the prompt.
Read the cost checksAI securityOpenAI Astra hit the critical cyber threshold. Now comes the ai model security trap
Why frontier cyber capability arrives with safeguards, and what they cost builders.
Read the Astra pieceAI consultancyThe Price Is Real. The Bill Is Somewhere Else
Model price cuts, promotional rates and the contract clause that outlives them.
Read the pricing pieceWant the bake-off run on your own work, not a leaderboard?
folkfox runs blind model comparisons on real client tasks, prices the finished output rather than the token, and hands you a routing plan you can keep.
Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.