Claude Sonnet 5.5 (claude-sonnet-5-5, released September 28, 2026) costs the same as Sonnet 5, $2 per million input tokens and $10 per million output, and uses the same tokenizer. It’s far stronger, and Anthropic says 30%+ faster. On the launch table, Terminal-Bench 4.0 rises from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5%, and OSWorld 2.1 from 57.0% to 80.1%. The catch is the API. Five request shapes that worked on Sonnet 5 now return a 400, and one change alters the response shape with no error. Verdict: upgrade, but fix those six things first.
Below: each breaking change with its exact error and before/after JSON, then a checklist. For specs, see what is Claude Sonnet 5.5; for the previous jump, Claude Sonnet 5 vs Sonnet 4.6. Apidog keeps old and new requests side by side while you test.
Sonnet 5 vs Sonnet 5.5 at a glance
| Claude Sonnet 5 | Claude Sonnet 5.5 | |
|---|---|---|
| Price per MTok (in / out / cache read) | $2 / $10 / $0.20 | $2 / $10 / $0.20 |
| Context and output | 1M context | 1M context, 128K output |
Accepted thinking.type |
adaptive, disabled |
adaptive, between_tools |
Default display |
omitted |
omitted |
| Effort | low to max |
Same levels, recalibrated; API default high |
| Minimum cacheable prompt | 1,024 tokens | 512 tokens |
| Per-message effort, mid-conversation system messages | No | Yes |
Forced tool_choice |
Supported | 400 error |
| Opus-style cyber safeguards | No | Yes; high-risk cyber falls back to Sonnet 5 |
| Thinking blocks | No conversation check | Bound to model, conversation and account |
| Terminal-Bench 4.0 | 10.3% | 70.6% |
| CursorBench 4.0 | 34.1% | 55.5% |
| FrontierCode 1.1 (Main) | 42.4% | 46.2% (max), 52.1% (xhigh) |
| GDPval-AA v2.1 (Elo) | 1449 | 1844 |
| OSWorld 2.1 (partial) | 57.0% | 80.1% |
| HLE (with tools) | 54.9% | 64.5% |
| Retirement | Still served as the cyber fallback | Not sooner than September 28, 2027 |
Anthropic ran Terminal-Bench, HLE and OSWorld; Cursor ran CursorBench, Cognition FrontierCode, and Artificial Analysis GDPval-AA. Artificial Analysis’s own Terminal-Bench run gives 63.6% vs 14.1%, so the gap holds. FrontierCode has two 5.5 numbers because at max it more often fanned out review subagents, and the benchmark penalizes out-of-scope edits. See Claude Sonnet 5.5 benchmarks.
The five breaking changes
Each returns a 400 invalid_request_error for code that runs fine on Sonnet 5.
1. thinking: disabled is gone; send between_tools
Sonnet 5.5 rejects thinking: {"type": "disabled"}:
"thinking.type.disabled" is not supported for this model. Use "thinking.type.between_tools" for the lowest thinking setting, or "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.
Send between_tools, the lowest setting. It skips up-front thinking; progress notes between tool calls still arrive as thinking blocks, which you pass back unchanged. It works only at low, medium or high effort, takes no other field (display, budget_tokens or block_binding return a 400), and locks effort for the conversation.
// Before (claude-sonnet-5)
{"model": "claude-sonnet-5", "max_tokens": 16000,
"thinking": {"type": "disabled"},
"output_config": {"effort": "xhigh"}}
// After (claude-sonnet-5-5)
{"model": "claude-sonnet-5-5", "max_tokens": 16000,
"thinking": {"type": "between_tools"},
"output_config": {"effort": "high"}}
If you need xhigh or max, omit thinking so adaptive thinking runs.
2. Forced tool_choice returns a 400
A tool_choice of any or tool fails, including on the token-counting endpoint:
tool_choice: type "tool" and "any" are not supported for this model.
Send auto, mark the tool strict: true (every object needs additionalProperties: false), and say in the prompt when to use it. The model can now reply in text instead, so handle turns with no tool call.
// Before (claude-sonnet-5)
"tool_choice": {"type": "tool", "name": "get_weather"}
// After (claude-sonnet-5-5)
"tools": [{"name": "get_weather",
"input_schema": {"type": "object",
"properties": {"location": {"type": "string"}},
"required": ["location"], "additionalProperties": false},
"strict": true}],
"tool_choice": {"type": "auto"},
"messages": [{"role": "user",
"content": "What's the weather in Paris? Use the get_weather tool."}]
The cap is 20 strict tools per request. On Amazon Bedrock, strict tools aren’t available for Sonnet 5.5: send auto without strict and validate the input in your code.
3. Thinking blocks are bound to the model, conversation and account
Each Sonnet 5.5 thinking block is signed over everything before it: system, tools and earlier messages. For accounts created on or after August 31, 2026 (00:00 UTC), replaying a block after an edit returns a 400 on the Claude API, Bedrock and Google Cloud:
messages.1.content.0: Invalid `signature` in `thinking` block. The block is bound to a different conversation. Remove the block, or set `thinking.block_binding.prefix_mismatch_behavior` to "drop_block".
Older accounts don’t enforce this by default, so a clean run on an old key proves nothing. Keep conversations append-only and change instructions or tools with mid-conversation system messages. If you must edit, send anthropic-beta: thinking-binding-controls-2026-08-01 and drop mismatched blocks:
"thinking": {"type": "adaptive",
"block_binding": {"prefix_mismatch_behavior": "drop_block"}}
That works only with adaptive thinking; with between_tools, strip thinking blocks from the edited turn on. Sonnet 5.5 reads blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier; it drops Opus 5, Opus 5.5, Fable and Mythos blocks, and Sonnet 5.5 blocks from another account, without failing the request. No other model reads its blocks. See Claude Fable 5.1 preserved thinking.
4. computer_20251124 fails on the Claude API and Google Cloud
There, computer use needs the new toolset. The error begins:
'claude-sonnet-5-5' does not support tool types: computer_20251124.
// Before (claude-sonnet-5)
"tools": [{"type": "computer_20251124", ...}]
// After (claude-sonnet-5-5, Claude API and Google Cloud)
"tools": [{"type": "computer_toolset_20260801"}]
Drop the old computer-use beta header and update your loop for member tool_use blocks, batched actions and toolset_name on results. Bedrock still accepts computer_20251124; computer_20250124 fails everywhere.
5. Some advisor pairings are rejected
With the advisor tool (beta), a Sonnet 5.5 executor accepts only Opus 5, Opus 5.5, Sonnet 5.5, Fable 5, Fable 5.1, Mythos 5 or Mythos 5.1 as advisor. Sonnet 5, Opus 4.8 and Opus 4.7 advisors now return a 400. Advice also arrives encrypted as an advisor_redacted_result block, so code that parses advice text gets nothing.
The silent change: text between tool calls moves into thinking blocks
This one fails nothing. On Sonnet 5, notes between tool calls came back as text. On Sonnet 5.5, anything longer than a sentence or two arrives as a progress-update thinking block, empty under the default display: "omitted". An agent UI that streams those notes goes quiet, with no error. Three fixes:
display: "updates"(adaptive thinking, beta headerthinking-display-updates-2026-08-18) returns the updates alone. Without the header, it’s rejected.display: "summarized"mixes the updates with reasoning summaries.between_toolsreturns the text with nodisplayneeded.
// Header: anthropic-beta: thinking-display-updates-2026-08-18
"thinking": {"type": "adaptive", "display": "updates"}
Behavior changes with no code change
These don’t fail requests, but they change output and cost. The prompting guide has fixes.
- Effort is recalibrated. Same names, different amounts of thinking. Re-run your sweep: start at
high,mediumfor well-specified agentic coding,mediumorlowfor chat; savexhighandmaxfor measured gains. - Changing top-level
effortbetween requests invalidates the prompt cache. Use the beta per-message effort instead. - It thinks before almost every reply from
mediumup. Lower the effort; prompting it to think less isn’t reliable. - It checks in early at
lowandmediumon long agentic tasks. - It adds unrequested tests, docs and files, plus reviewer subagents at
xhighandmax. Anthropic’s suggested prompt cut session cost atmaxby about a third. - More refusals. It’s the first Sonnet with Opus-style cyber safeguards. A decline is HTTP 200 with
stop_reason: "refusal"and astop_detailscategory; beta server-side fallback on the Claude API (fallbacks: "default") retriescyberandfrontier_llmdeclines on Sonnet 5. - Sampling parameters. Non-default
temperature,top_portop_kreturn a 400. Anthropic lists this under moves from Sonnet 4.6 and earlier, so you’ve likely removed them already.
Cost per task: same price, fewer dollars per result
Anthropic’s launch post says Sonnet 5.5 “costs up to 30% less per task than its predecessor” in its testing. Prices are identical, so the saving comes from fewer tokens and steps. Its per-effort charts show lower effort on 5.5 beating Sonnet 5’s best run:
| Benchmark (Anthropic’s charts) | Sonnet 5.5 | Sonnet 5, best run |
|---|---|---|
| Terminal-Bench 4.0 | 28.8% at medium, $0.83 | 10.3% at max, $11.62 |
| FrontierCode 1.1 | 49.4% at high, $0.42 | 42.7% at xhigh, $10.07 |
| CursorBench 4.0 | 35.8% at low, $0.50 | 34.1% at max, $7.17 |
CursorBench costs are Anthropic’s estimates at list prices. The flip side is max: per OfficeChai’s write-up of Artificial Analysis data, Sonnet 5.5 at max uses about 193,000 output tokens per index task, and cost per task runs about 50% above Sonnet 5. The savings live at high and below; see Claude Sonnet 5.5 pricing.
Migration checklist
Swap claude-sonnet-5 for claude-sonnet-5-5, then run the six checks from Anthropic’s migration guide:
- Replace
disabledwithbetween_tools, athigheffort or below. - Replace forced
tool_choicewithauto,strict: trueand a prompt line (on Bedrock, validate in code). - Keep history append-only; use mid-conversation system messages for changes.
- Move computer use to
computer_toolset_20260801on the Claude API and Google Cloud. - Pick a supported advisor, and stop parsing advice text.
- Set
thinking.displayif your UI shows text between tool calls.
Then re-run your effort sweep. Claude Code can automate the migration:
/claude-api migrate this project to claude-sonnet-5-5
The Opus equivalent is Claude Opus 5.5 vs Opus 5 migration.
Turn the migration into a regression test in Apidog
In Apidog, save three requests to https://api.anthropic.com/v1/messages in one project, all on the same environment:

- Baseline: your current Sonnet 5 body.
- Old body, new model: only
modelchanged toclaude-sonnet-5-5. Assert status 400 and an error message mentioningbetween_toolsortool_choice. - Migrated: the fixed body. Assert status 200, a
stop_reasonthat isn’trefusalormax_tokens, and atool_useblock for tool requests.
Keep ANTHROPIC_API_KEY in the environment and reference it as {{ANTHROPIC_API_KEY}} in the x-api-key header. Compare usage.output_tokens across effort levels for your own cost-per-task number. Saved as a test scenario, a regression to disabled fails the run instead of production. Request basics: how to use the Claude Sonnet 5.5 API.
FAQ
Is Claude Sonnet 5.5 more expensive than Sonnet 5? No. Both cost $2/$10 per million tokens with $0.20 cache reads, and the tokenizer is the same.
Why do I get “thinking.type.disabled is not supported” on Sonnet 5.5? disabled was removed. Send thinking: {"type": "between_tools"} at low, medium or high effort, with no other thinking fields.
Do Sonnet 5 conversations carry over to Sonnet 5.5? Yes. Sonnet 5.5 reads Sonnet 5 thinking blocks. Moving back loses them: no other model reads Sonnet 5.5 blocks.
Should I move everything to Sonnet 5.5? For most workloads, yes. Watch cyber-adjacent work, which may fall back to Sonnet 5, and Bedrock tool calls, which lose strict mode. Our Claude Sonnet 5 guide covers the model you’re leaving.
Your next step
Save your Sonnet 5 request and its 5.5 twin side by side, confirm the 400, fix it, and move traffic once the fix passes. Download Apidog to build that test.



