What Shipped
Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family after Opus 5.5. Model ID is claude-sonnet-5-5.
The short version: same price per token, meaningfully faster, and close enough to Opus 5.5 on several benchmarks that the price gap becomes the deciding factor for a lot of work.
Pricing: Unchanged
- Input: $2 per million tokens
- Output: $10 per million tokens
- Cache reads: $0.20 per million
- Cache writes: $2.50 per million
Those are identical to Sonnet 5. Nothing got cheaper per token.
So Where Does "Up to 30% Less" Come From?
Efficiency, not price. Sonnet 5.5 uses fewer tokens to finish the same task, and it batches tool calls more often than Sonnet 5 did, which cuts the number of round trips a job needs.
This is worth understanding properly because it changes who benefits. If your workload is a single prompt returning a paragraph, your bill will look much the same. If it is an agent making twenty tool calls to complete something, the batching is where the saving lives. The gain scales with how agentic your usage is.
Output generation is separately more than 30% faster than Sonnet 5.
The Benchmarks
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% |
| GDPval-AA v2.1 | 1,844 | 1,449 | 1,846 |
| Chartography | 61.6% | 15.6% | 64.4% |
Also published: FrontierCode 1.1 at 46.2% on Max effort and 52.1% on Xhigh, Humanity's Last Exam at 64.5% with tools, and OSWorld 2.1 at 80.1% partial.
The line that will get quoted everywhere is Terminal-Bench: Sonnet 5.5 at 70.6% against Opus 5.5 at 66.4%. The mid-tier model beating the flagship on agentic terminal work is a real result and it is the most interesting number here.
The Number That Is Not What It Looks Like
Terminal-Bench 4.0 going from 10.3% to 70.6% looks like a sevenfold capability jump. It mostly is not.
Sonnet 5 was hitting timeouts and token limits on that benchmark, which is why its score was so low. A model that runs out of budget mid-task scores near zero regardless of whether it knew what to do. So a large part of that gap measures Sonnet 5.5 finishing, not Sonnet 5.5 being seven times smarter.
The same caution applies to Chartography at 15.6% to 61.6%.
The honest comparisons in that table are the ones where Sonnet 5 was already functioning: CursorBench 34.1% to 55.5%, and GDPval-AA 1,449 to 1,844. Those are solid, real improvements. They are just not sevenfold.
It is also worth stating plainly that all of these figures are Anthropic's own. No independent third-party verification has been published yet.
Where Opus 5.5 Still Wins
Anthropic says it directly: Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.
The benchmark table supports that. Sonnet 5.5 edges ahead on Terminal-Bench, but trails on CursorBench (55.5% vs 57.8%), Chartography (61.6% vs 64.4%) and is a hair behind on GDPval-AA (1,844 vs 1,846).
The pattern: Sonnet 5.5 has closed the gap on well-scoped, mechanical, high-volume work. Anthropic positions it for everyday tasks, bug fixing, and producing documents, slides and spreadsheets. Long-horizon reasoning where the task shape is not clear at the outset is still Opus territory.
The Part Almost Nobody Is Covering
Two safeguards shipped with this model that have never been on a Sonnet before.
Cyber restrictions
Sonnet 5.5's cybersecurity capability is comparable to Opus 5, so it launches with restrictions previously reserved for Anthropic's most capable models. Higher-risk cyber requests visibly fall back to Sonnet 5.
That lands in a specific week. OpenAI has just declared its Astra model the first to meet the Critical cybersecurity threshold in its Preparedness Framework, and Google shipped Gemini 3.8 Flash Cyber through a gated defender programme. Three labs, three capability-gated cyber releases, within days of each other and days before the White House AI meeting. We covered the Astra side of that here.
Anti-distillation
Sonnet 5.5 is the first Sonnet with classifiers that block reasoning extraction, tying reasoning to specific user accounts. The trigger was researchers decoding more than 315,000 thinking blocks from public traces.
If you build anything that reads or logs Claude's reasoning output, check your pipeline against this before you upgrade.
Should You Switch?
- On Sonnet 5 already: yes. Same price, faster, better. The only real risk is the cyber fallback and the reasoning-extraction classifiers if either touches your use case.
- On Opus 5.5 for high-volume mechanical work: test Sonnet 5.5 against your own tasks. On benchmarks the gap is now small, and the cost difference is not.
- On Opus 5.5 for open-ended judgment work: stay. Anthropic is telling you to, and that is not the kind of thing a vendor says lightly about its cheaper model.
- Running agents with many tool calls: this is the workload the efficiency gains were built for. Measure your own token count before and after rather than trusting the 30%.
Where to Get It
Claude Platform, AWS, Google Cloud and Microsoft Azure, plus all the consumer Claude apps. A zero data retention option is available.
For what the free tier covers and where the API limits sit, see our breakdown of Claude free plan limits and API credits. If you are running Claude in agent workflows, our guide to Claude Managed Agents covers the deployment side, and Claude computer use on Windows covers the desktop path - relevant here given the OSWorld 2.1 result.
FAQ
Is Claude Sonnet 5.5 more expensive than Sonnet 5?
No. Pricing is identical: $2 per million input tokens, $10 per million output, $0.20 per million cache reads.
How is it 30% cheaper if the price is the same?
Cost per task falls because the model uses fewer tokens and batches tool calls, not because the rate changed. Simple single-turn prompts will see less benefit than multi-step agent workloads.
Is Sonnet 5.5 better than Opus 5.5?
On Terminal-Bench 4.0, yes - 70.6% against 66.4%. On CursorBench, Chartography and GDPval-AA, Opus 5.5 is still ahead, though narrowly. Anthropic says Opus remains clearly stronger on complex open-ended work.
Did Sonnet 5 really only score 10.3% on Terminal-Bench?
Yes, but that figure reflects timeouts and token limits rather than pure capability, so the improvement is smaller than the raw numbers suggest.
What is the model ID?
claude-sonnet-5-5
Are there new restrictions?
Yes. Higher-risk cyber requests fall back to Sonnet 5, and classifiers now block reasoning extraction by tying reasoning to user accounts. Both are firsts for a Sonnet model.