

Claude Haiku 5.5 is 90% cheaper per token, but the bill depends on four other numbers
Anthropic released Claude Haiku 5.5 on 7 October 2026 with the lowest per-token prices in its Haiku line. The headline cut is real, but it applies to one price band. The prowl for a cheaper bill starts with the token count, the effort setting, the cache and the migration checklist.
By Katie Delaney / 2026-10-07 / 11 min read

Claude Haiku 5.5 has one price band for the 90% cut#
Claude Haiku 5.5 is listed on Anthropic's pricing page at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. The same page lists Claude Haiku 4.5 at $1 and $5. That is the 90 per cent in the headline, and it applies to one band.
Above 100,000 tokens the rate is $0.50 input and $2.50 output, half the Haiku 4.5 rate. Anthropic's launch gives a second figure that answers a different question. Across its own workloads, the company says the model costs around 75 per cent less to run on average. The first figure is a price. The second is a bill, and a bill depends on prompt length, token counts and the effort you choose.
Batch processing is cheaper again for work that can wait: $0.05 input and $0.25 output per million tokens for prompts up to 100,000 tokens. Anthropic's overview lists the batch discount as 50 per cent on input and output.
| Item | Value |
|---|---|
| Haiku 5.5, up to 100K | $0.1 |
| Haiku 5.5, over 100K | $0.5 |
| Haiku 4.5 | $1 |
The tokeniser changes the count#
The second lever is the tokeniser. Anthropic's overview says the same text counts as approximately 30% more tokens on Claude Haiku 5.5 than on Haiku 4.5, because the model uses the newer tokeniser that Claude 4.7 and later models share. The exact increase depends on the content. Read the price on its own and the trail looks straight. The number of tokens each prompt produces is the other lever, and it moved up.
We ran the arithmetic on a plain, illustrative workload: 100,000 calls a month, each with 1,000 input tokens and 100 output tokens counted on the Haiku 4.5 tokeniser. On Haiku 4.5 prices that is $150. The same prompts produce 30 per cent more tokens on the newer tokeniser, so the same text would cost $195 at Haiku 4.5 prices. Move those uplifted tokens onto Claude Haiku 5.5 prices, which means 130 million input tokens at $0.10 and 13 million output tokens at $0.50 per million, and the bill falls to $19.50. Against the $150 starting point, $19.50 is a saving of 87 per cent. With no uplift at all, the price cut alone would give $15, or 90 per cent, and the three-point gap is the cost of the tokeniser.
This is our calculation from published prices and Anthropic's approximate uplift. It is not a measurement of any live workload, and the exact uplift depends on the content, as the documentation says. Your own sample will decide the number.
| Item | Value |
|---|---|
| Haiku 4.5 bill | +150 (running total 150) |
| Tokeniser uplift | +45 (running total 195) |
| Lower price | -175.5 (running total 19.5) |
| Haiku 5.5 bill | 19.5 |

The benchmarks are vendor numbers#
Anthropic's launch table compares Claude Haiku 5.5 with Haiku 4.5, GPT-6 Luna and Sonnet 5.5. The biggest move is computer use: 72.4 per cent on OSWorld 2.1's offline subset, against 15.7 per cent for Haiku 4.5. Terminal-Bench 4.0 goes from 0.0 to 39.2 per cent. The announcement is the source for every row.
The table is thin on method. Only one row names an effort level, and the documented default is medium. Anthropic's effort guide lists medium as the default on the Claude API and in Claude Code. Read the scores as the vendor's measurements until your own scored set says otherwise.
An independent index gives a more cautious read. Artificial Analysis scored Haiku 5.5 at 34 on its Intelligence Index at medium effort, ranked 15th of 182 models when we checked. It published no speed figures for the model at that point.
The headline computer-use score also needs a footnote. Kingy.ai, which analysed launch-day evaluations without running its own matched test, notes that the same evaluation reports a 37.1% rate for completing every checkpoint, so the 72.4% figure is partial credit.
| Item | Value |
|---|---|
| OSWorld 2.1 offline | 15.7 to 72.4 |
| Last Exam, no tools | 10.2 to 45.9 |
| Terminal-Bench 4.0 | 0 to 39.2 |
| Chartography, no tools | 6.4 to 46.4 |
Switching is a migration, not a model string#
A clean Claude Haiku 5.5 migration starts with Anthropic's migration guide, not the model name. The checklist has ten items for code that calls Haiku 4.5, and several of them fail on the first request. The What's new page lists the same changes by type.
| Change | What happens | What to do |
|---|---|---|
| budget_tokens in manual extended thinking | Returns a 400 error | Use adaptive thinking, with effort as the dial |
| temperature, top_p and top_k | Non-default values return a 400 error | Omit them, or keep the defaults the guide allows |
| Assistant message prefill | Rejected | End messages with a user turn and move the instruction into the prompt |
| Computer use on the Claude API or Google Cloud | The old computer_20250124 tool returns a 400 error | Move to computer_toolset_20260801 |
| Edited earlier turns when thinking blocks are sent back | Returns a 400 error | Keep conversations append-only |
| Thinking blocks replayed through another account | Dropped before the model sees them | Replay each conversation through the account that produced it |
| stop_reason refusal | Safety classifiers can decline, with no fallback | Branch on the stop reason rather than retrying blindly |
| Priority Tier commitments made for Haiku 4.5 | Not supported on Haiku 5.5 | Plan that capacity separately |
- budget_tokens in manual extended thinkingReturns a 400 errorUse adaptive thinking, with effort as the dial
- temperature, top_p and top_kNon-default values return a 400 errorOmit them, or keep the defaults the guide allows
- Assistant message prefillRejectedEnd messages with a user turn and move the instruction into the prompt
- Computer use on the Claude API or Google CloudThe old computer_20250124 tool returns a 400 errorMove to computer_toolset_20260801
- Edited earlier turns when thinking blocks are sent backReturns a 400 errorKeep conversations append-only
- Thinking blocks replayed through another accountDropped before the model sees themReplay each conversation through the account that produced it
- stop_reason refusalSafety classifiers can decline, with no fallbackBranch on the stop reason rather than retrying blindly
- Priority Tier commitments made for Haiku 4.5Not supported on Haiku 5.5Plan that capacity separately
Two smaller traps catch teams out. Thinking tokens count toward max_tokens, so a small cap can end inside a thinking block before any text appears. And changing the top-level effort between requests invalidates prompt caching, so hold effort constant inside a cached conversation, as the effort guide explains. The cap is the trap people hit last, hiding in the brush.

Effort is the new dial, and the bigger models keep the hard work#
Claude Haiku 5.5 is the first Haiku-class model with an adjustable effort setting. Anthropic says this lets users decide whether to optimise for cost or intelligence. The effort guide supports levels from low to max. Medium is the default, and low is the suggested setting for simple, high-volume work and subagents. The Verge reported the company's claim that the model is its fastest to date.
Much of the early conversation on X has stayed on the price. One practitioner put the right job more usefully than the price table:
I'd try it for turning CI logs into a short failure summary, then leave the fix to a bigger model.
That is close to the shape Anthropic gives the model: high-volume, cost-sensitive tasks such as classification, routing and extraction, and subagent work. Narrow, repeatable and checkable is the test. If a person can say in a few seconds whether the answer is right, a small model such as Claude Haiku 5.5 is a sensible first pass.
The bigger models keep the hard work. Anthropic's announcement says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding, and it describes Haiku as a subagent that pairs well with them.
Refusals need a plan too. Anthropic says Haiku 5.5 runs safety classifiers that can decline a request, and it has no server-side fallback. The migration guide tells developers to handle the refusal stop reason, so route those cases to a person or to a different approved model.
Where the savings land in marketing work#
Small AI model costs only make sense per correct answer. Most growth teams run tagging, classification, field extraction and first-pass summaries across search-term lists, competitor pages, review text and inbound messages. Those calls multiply. The trail is short, because most teams already know which calls they run most.
Three jobs fit the description Anthropic gives and need no leap of faith. Group inbound search terms into intent buckets before a person checks them. Pull named fields, such as price, plan and claim, from pages you already hold. Route customer messages to the right queue. Each has an answer someone can check, and that check is what makes Claude Haiku 5.5 safe at volume.
Two limits apply to any model. Anything that reaches a regulator, a client or a patient needs human review whatever the price. And a cheap model on the wrong task costs less per call and more per error. Price the errors as well as the tokens.
Launch coverage puts GPT-6 Luna at the same $0.10 and $0.50 per million tokens for short prompts, so the headline price alone does not settle the choice. VentureBeat reported the match. Quality, latency and reliability are separate questions, and the vendor table does not answer them for your workload.
Claude Haiku 5.5 is a real price cut on the calls that multiply, and the quarry is the answer, not the invoice. Teams that keep the saving will be the ones that measure answers as carefully as the bill. That is the work at the core of folkfox's AI consultancy. Paid media teams can start with our PPC work, content teams with content marketing, and search teams with SEO and GEO. To talk it through, use the contact page.
Frequently asked questions#
Is Claude Haiku 5.5 cheaper than Claude Haiku 4.5?
Yes, for prompts up to 100,000 tokens. Anthropic's pricing page lists $0.10 per million input tokens and $0.50 per million output tokens, against $1 and $5 for Claude Haiku 4.5. The newer tokeniser counts about 30 per cent more tokens for the same text, so the saving on an illustrative workload is about 87 per cent rather than 90.
What is the Haiku 5.5 pricing for long prompts?
Prompts over 100,000 tokens cost $0.50 per million input tokens and $2.50 per million output tokens, which is half the Haiku 4.5 rate in both directions. Batch processing is cheaper again for prompts up to 100,000 tokens, at $0.05 input and $0.25 output, according to the pricing page.
Will my Haiku 4.5 code work on Haiku 5.5 without changes?
Not always. Manual extended thinking with budget_tokens returns an error, and so do non-default temperature, top_p and top_k values. Assistant message prefill is rejected, and computer use moves to a new toolset. Work through the migration guide's checklist before you change the model string in production.
What does the Haiku 5.5 effort setting do?
Effort trades response quality against speed and cost. Medium is the default on the Claude API and in Claude Code, and adaptive thinking is on from the start. Thinking tokens count toward max_tokens, so leave room for them when you set the cap.
When should a team use a bigger Claude model instead?
Anthropic says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks. Keep them for multi-step work where an error costs more than the tokens, and use Claude Haiku 5.5 for narrow, repeatable calls that a person can check quickly.
Read more on this topic#
OpenAI API rate limits drop to three tiers
The same question about model tiers, from the OpenAI side, on the morning edition.
Read the field noteAI consultancyWhat an AI subscription is really worth: SemiAnalysis measured the meter, not the bill
Why a headline price and a monthly bill are different questions.
Read the field noteAI consultancyOpen weight models meet a 3,800-GPU reality. Mistral's preview exposes the infrastructure underneath
What cheap model access does and does not change in your supply chain.
Read the field noteWant a cost model before you switch?
folkfox builds the test set, the cost-per-correct-answer sheet and the routing rules for AI consultancy clients. Tell us the workload and we will show you where the savings are real.
Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.
