Claude Sonnet 5.5 keeps the price, promises fewer tokens, and breaks forced tool calls
Anthropic's new mid-tier model costs the same as Sonnet 5 and claims to finish tasks faster. The catch is in the docs: five breaking API changes.
Anthropic released Claude Sonnet 5.5 on Monday, September 28, as the second model in its 5.5 line after Opus 5.5. The price did not move: $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5, with a 1M-token context window and 128K tokens of output. What Anthropic is selling is efficiency. In its own testing, the company says Sonnet 5.5 generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task, because it needs fewer tokens and tool calls to finish the same work (Anthropic).
For developers the more consequential part is in the documentation, not the launch page. Anthropic lists five breaking changes for code that already runs on Sonnet 5 (release notes). Swapping the model ID is not a safe upgrade on its own.
What Anthropic claims
All the benchmark figures below come from Anthropic's announcement; no independent evaluation had been published when this ran. By the company's numbers, Sonnet 5.5 lands close to Opus 5.5 on several tests while costing half as much per token ($2/$10 versus $4/$20):
- Terminal-Bench 4.0 (multi-step tasks in a command line): 70.6%, versus 66.4% for Opus 5.5 at its highest effort setting.
- CursorBench 4.0: 55.5%, versus 34.1% for Sonnet 5 and 57.8% for Opus 5.5.
- FrontierCode 1.1: 46.2%, versus 42.4% for Sonnet 5, 54.4% for Opus 5.5 and 49.3% for OpenAI's GPT-6 Sol.
- OSWorld 2.1 (computer use, partial credit): 80.1%, versus 57.0% for Sonnet 5 and 81.8% for Opus 5.5.
- GDPval-AA v2.1 (knowledge work, run by Artificial Analysis): 1844, versus 1449 for Sonnet 5 and 1846 for Opus 5.5.
The footnotes matter. Anthropic says Artificial Analysis tested a pre-release deployment that had a bug degrading structured-output responses, since fixed. It also says Sonnet 5.5 scored lower on FrontierCode at Max effort than at Xhigh, because at Max it more often launched a multi-agent code review that timed out or made edits beyond the task's scope. Anthropic itself says Opus 5.5 remains clearly stronger on complex, open-ended work that needs sustained judgment.
TechCrunch summarized the launch as a speed play and reported that Anthropic plans a new Haiku model in the coming weeks, without a firm date.
The five breaking changes
According to Anthropic's migration notes, each of these returns an error or silently changes behavior for code written against Sonnet 5:
- No more thinking: disabled. Sending it returns a 400. The lowest setting is now between_tools, which turns off up-front thinking and only works at high effort or below. Manual thinking budgets also return a 400.
- No forced tool use. tool_choice of type any or tool returns a 400. Only auto and none are accepted. Anthropic's recommended replacement is strict tool use or structured outputs, plus saying in the prompt when a tool applies.
- Thinking blocks are bound to the model and conversation. Sonnet 5.5 blocks can't be read by any other model, and for accounts created on or after August 31, 2026, replaying a block after editing the system prompt, tools or an earlier message returns a 400 by default.
- The older computer_20251124 tool is rejected on the Claude API and Google Cloud; computer use there requires the newer computer_toolset_20260801. Amazon Bedrock still accepts the old tool.
- Some advisor-tool pairings fail: Opus 4.8, Opus 4.7 and Sonnet 5 are rejected as advisors to a Sonnet 5.5 executor.
There is also a quieter change. Between tool calls, longer progress notes now come back as thinking blocks, which are empty by default. An app that streams those notes to its users will simply go quiet mid-task, with no error, until it changes the display setting or uses between_tools.
What changes for prompts and model choice
The forced-tool-use removal is the one most likely to break production code. A common pattern for getting JSON out of a model is to define a single tool and force the model to call it. That pattern returns a 400 on Sonnet 5.5. Teams that rely on it should move to structured outputs or strict tool use before switching, and re-test any prompt that assumed the model had no choice.
Effort also needs re-tuning. Anthropic says effort levels are recalibrated, so the same setting does not produce the same amount of thinking as on Sonnet 5. Its guidance is to start at high by default, drop to medium for well-specified agentic coding, and use medium or low for chat and latency-sensitive work. Default effort on the API is high.
On cost, the per-token price is unchanged, so any saving depends on the model actually using fewer tokens on your workload. The minimum cacheable prompt drops from 1,024 tokens on Sonnet 5 to 512, which can help prompt-caching on shorter system prompts. Non-default temperature, top_p or top_k still return a 400.
Safeguards
Sonnet 5.5 is the first Sonnet to ship with the cyber safeguards Anthropic uses on its most capable models. Per the announcement, higher-risk cybersecurity requests visibly fall back to Sonnet 5, and security teams can apply to an expanded Cyber Verification Program for broader access. The model also adds classifiers against extracting its internal reasoning, and its thinking blocks only work in the account that produced them.
Availability
The model ID is claude-sonnet-5-5 on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. AWS confirmed availability on Bedrock and AWS GovCloud (US) the same day. Sonnet 5 is now listed as legacy; Sonnet 5.5's retirement date is set no sooner than September 28, 2027 (model page).
Why this matters
- Sonnet is the default workhorse for many API and coding-agent workloads, so a same-price upgrade with claimed per-task savings affects a lot of production traffic.
- Code that forces a tool call to get JSON, or disables thinking, will return 400 errors on Sonnet 5.5 and must be changed before switching.
- All benchmark and cost figures are Anthropic's own; independent evaluation is not yet available.
Key takeaways
- Price unchanged at $2 input / $10 output per million tokens, 1M context, 128K output.
- Anthropic claims 30%+ faster output and up to 30% lower cost per task versus Sonnet 5.
- tool_choice any/tool and thinking: disabled now return 400; use structured outputs or strict tool use, and between_tools.
- Effort levels are recalibrated; re-run effort sweeps rather than copying Sonnet 5 settings.
Sources
- AnthropicPrimaryIntroducing Claude Sonnet 5.5anthropic.com
- Anthropic (Claude Platform Docs)PrimaryWhat's new in Claude Sonnet 5.5platform.claude.com
- Anthropic (Claude Platform Docs)PrimaryClaude Platform release notes (September 28, 2026)platform.claude.com
- Anthropic (Claude Platform Docs)PrimaryClaude Sonnet 5.5 model pageplatform.claude.com
- Amazon Web ServicesPrimaryClaude Sonnet 5.5 now available on AWSaws.amazon.com
- TechCrunchAnthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partnertechcrunch.com
- claude
- sonnet
- api
- tool-use
- structured-outputs
- migration
- pricing
- Anthropic
- Amazon Web Services
- Claude Sonnet 5.5
- Claude Sonnet 5
- Claude Opus 5.5