THE SHORT VERSION
● Under 100K tokens: $0.10 in / $0.50 out per million — a 90% cut
● Over 100K tokens: $0.50 in / $2.50 out — a 50% cut
● Anthropic's own estimate: about 75% lower on realistic workloads
● Biggest jump: OSWorld 2.1 from 15.7% to 72.4%
● Read the asterisk: Terminal-Bench 39.2% is at max effort. Default is about 20%
Anthropic shipped Claude Haiku 5.5 on 7 October 2026. The headline everywhere is a 90% price cut. That number is real but it applies to one of two tiers, and the gap between them is where your bill actually gets decided.
The pricing, all of it
| Per 1M tokens |
Under 100K |
Over 100K |
Haiku 4.5 |
| Input |
$0.10 |
$0.50 |
$1.00 |
| Output |
$0.50 |
$2.50 |
$5.00 |
| Cache reads |
$0.01 |
$0.05 |
$0.10 |
| Cache writes |
$0.125 |
$0.625 |
$1.25 |
The threshold is the prompt size, and crossing it multiplies everything by five. A retrieval pipeline that stuffs context, a long chat history, a document summariser — these sit above 100K routinely. If that is your workload, your cut is 50%, and the 90% headline was never yours.
Anthropic's own figure is the honest one: about 75% lower workload cost, which accounts both for the mix of request sizes and for a tokenizer that produces somewhat more tokens than the previous one. Budget against 75%, not 90%.
What it costs against everything else
| Model |
Input / 1M |
Output / 1M |
| Claude Haiku 5.5 |
$0.10 |
$0.50 |
| GPT-6 Luna |
$0.10 |
$0.50 |
| Gemini 3.5 Flash-Lite |
$0.30 |
$2.50 |
| Gemini 3.8 Flash |
$0.75 |
$3.75 |
| Grok 4.3 |
$1.25 |
$2.50 |
| Claude Sonnet 5.5 |
$2.00 |
$10.00 |
Haiku 5.5 and GPT-6 Luna are priced identically to the cent. That is not coincidence — it is a floor being set. Gemini 3.8 Flash is seven times the output price and is on a promotional rate until 31 December 2026, so expect movement there.
The benchmarks, including the one with a footnote
| Benchmark |
Haiku 5.5 |
Haiku 4.5 |
GPT-6 Luna |
Sonnet 5.5 |
| GDPval-AA v2.1 |
1,620 |
735 |
1,437 |
1,840 |
| AA-Briefcase v1.1 |
1,578 |
614 |
1,336 |
1,824 |
| OSWorld 2.1 (offline) |
72.4% |
15.7% |
48.9% |
83.9% |
| Terminal-Bench 4.0 |
39.2%* |
0.0% |
16.4% |
70.6% |
| FrontierCode 1.1 |
46.4% |
— |
42.4% |
52.1% |
* THE ASTERISK THAT MATTERS
Terminal-Bench 39.2% is measured at maximum effort. At medium effort, which is the default, it scores about 20%.
If you deploy on defaults and expect the published number, you get roughly half of it. These are also vendor-reported results.
Set the asterisk aside and the generational jump is genuine. Haiku 4.5 scored 0.0% on Terminal-Bench 4.0 — not a weak score, a complete failure. OSWorld went from 15.7% to 72.4%. On knowledge work, Haiku 5.5 now beats GPT-6 Luna at identical pricing on every line in that table.
Sonnet 5.5 still wins everywhere, which is the point of having two models. The question is no longer whether Haiku can do the job — it is which jobs now stop needing Sonnet.
The quieter change
Anthropic also cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens. No announcement fanfare, and for anyone running a large cached system prompt across high request volume it is worth more than the Haiku news.
Who should switch
PRACTICAL READ
● Switch now: classification, extraction, routing, short-prompt high-volume work
● Check your token sizes first: anything regularly over 100K gets the 50% tier
● Test at your actual effort setting: not at max, where the benchmarks were run
● Stay on Sonnet for: long-horizon agent work. 39% versus 70% on Terminal-Bench is the gap
Availability is immediate through Anthropic's platform, AWS, Google Cloud and Microsoft Azure. The API identifier is claude-haiku-5-5.
If you are weighing paid API work against what is free, our breakdown of Claude free plan limits and API credits covers where the no-cost ceiling sits.
FAQ
How much does Claude Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 output for prompts under 100,000 tokens. Above that, $0.50 input and $2.50 output.
Is it really 90% cheaper?
For prompts under 100,000 tokens, yes, against Haiku 4.5. Above that it is 50%. Anthropic's own realistic workload estimate is about 75%, which also accounts for a tokenizer that produces more tokens than before.
Haiku 5.5 or GPT-6 Luna?
Identical pricing at $0.10 / $0.50. On the published benchmarks Haiku 5.5 leads on all of them — 72.4% versus 48.9% on OSWorld, 39.2% versus 16.4% on Terminal-Bench, 1,620 versus 1,437 on GDPval. Both are vendor-reported, so test on your own workload.
Can it replace Sonnet for agent work?
Not for long-horizon agents. Terminal-Bench 4.0 is 39.2% at max effort against Sonnet's 70.6%, and about 20% at the default medium effort. For short tool-calling loops it is now viable where Haiku 4.5, at 0.0%, was not.
Did anything else change?
Yes. Sonnet 5.5 cache reads dropped from $0.20 to $0.10 per million tokens, which matters more than the Haiku change if you run a large cached system prompt at volume.
Sources