Claude Sonnet 5.5 vs Sonnet 5: What Changed, and the Breaking Changes to Fix Before You Switch

Sonnet 5.5 vs Sonnet 5: same $2/$10 price, far higher scores, and five breaking changes that return 400s. Exact errors and before/after JSON fixes.

Medy Evrard

29 September 2026

Claude Sonnet 5.5 vs Sonnet 5: What Changed, and the Breaking Changes to Fix Before You Switch

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Claude Sonnet 5.5 (claude-sonnet-5-5, released September 28, 2026) costs the same as Sonnet 5, $2 per million input tokens and $10 per million output, and uses the same tokenizer. It’s far stronger, and Anthropic says 30%+ faster. On the launch table, Terminal-Bench 4.0 rises from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5%, and OSWorld 2.1 from 57.0% to 80.1%. The catch is the API. Five request shapes that worked on Sonnet 5 now return a 400, and one change alters the response shape with no error. Verdict: upgrade, but fix those six things first.

Below: each breaking change with its exact error and before/after JSON, then a checklist. For specs, see what is Claude Sonnet 5.5; for the previous jump, Claude Sonnet 5 vs Sonnet 4.6. Apidog keeps old and new requests side by side while you test.

Sonnet 5 vs Sonnet 5.5 at a glance

Claude Sonnet 5 Claude Sonnet 5.5
Price per MTok (in / out / cache read) $2 / $10 / $0.20 $2 / $10 / $0.20
Context and output 1M context 1M context, 128K output
Accepted thinking.type adaptive, disabled adaptive, between_tools
Default display omitted omitted
Effort low to max Same levels, recalibrated; API default high
Minimum cacheable prompt 1,024 tokens 512 tokens
Per-message effort, mid-conversation system messages No Yes
Forced tool_choice Supported 400 error
Opus-style cyber safeguards No Yes; high-risk cyber falls back to Sonnet 5
Thinking blocks No conversation check Bound to model, conversation and account
Terminal-Bench 4.0 10.3% 70.6%
CursorBench 4.0 34.1% 55.5%
FrontierCode 1.1 (Main) 42.4% 46.2% (max), 52.1% (xhigh)
GDPval-AA v2.1 (Elo) 1449 1844
OSWorld 2.1 (partial) 57.0% 80.1%
HLE (with tools) 54.9% 64.5%
Retirement Still served as the cyber fallback Not sooner than September 28, 2027

Anthropic ran Terminal-Bench, HLE and OSWorld; Cursor ran CursorBench, Cognition FrontierCode, and Artificial Analysis GDPval-AA. Artificial Analysis’s own Terminal-Bench run gives 63.6% vs 14.1%, so the gap holds. FrontierCode has two 5.5 numbers because at max it more often fanned out review subagents, and the benchmark penalizes out-of-scope edits. See Claude Sonnet 5.5 benchmarks.

The five breaking changes

Each returns a 400 invalid_request_error for code that runs fine on Sonnet 5.

1. thinking: disabled is gone; send between_tools

Sonnet 5.5 rejects thinking: {"type": "disabled"}:

"thinking.type.disabled" is not supported for this model. Use "thinking.type.between_tools" for the lowest thinking setting, or "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.

Send between_tools, the lowest setting. It skips up-front thinking; progress notes between tool calls still arrive as thinking blocks, which you pass back unchanged. It works only at low, medium or high effort, takes no other field (display, budget_tokens or block_binding return a 400), and locks effort for the conversation.

// Before (claude-sonnet-5)
{"model": "claude-sonnet-5", "max_tokens": 16000,
 "thinking": {"type": "disabled"},
 "output_config": {"effort": "xhigh"}}

// After (claude-sonnet-5-5)
{"model": "claude-sonnet-5-5", "max_tokens": 16000,
 "thinking": {"type": "between_tools"},
 "output_config": {"effort": "high"}}

If you need xhigh or max, omit thinking so adaptive thinking runs.

2. Forced tool_choice returns a 400

A tool_choice of any or tool fails, including on the token-counting endpoint:

tool_choice: type "tool" and "any" are not supported for this model.

Send auto, mark the tool strict: true (every object needs additionalProperties: false), and say in the prompt when to use it. The model can now reply in text instead, so handle turns with no tool call.

// Before (claude-sonnet-5)
"tool_choice": {"type": "tool", "name": "get_weather"}

// After (claude-sonnet-5-5)
"tools": [{"name": "get_weather",
  "input_schema": {"type": "object",
    "properties": {"location": {"type": "string"}},
    "required": ["location"], "additionalProperties": false},
  "strict": true}],
"tool_choice": {"type": "auto"},
"messages": [{"role": "user",
  "content": "What's the weather in Paris? Use the get_weather tool."}]

The cap is 20 strict tools per request. On Amazon Bedrock, strict tools aren’t available for Sonnet 5.5: send auto without strict and validate the input in your code.

3. Thinking blocks are bound to the model, conversation and account

Each Sonnet 5.5 thinking block is signed over everything before it: system, tools and earlier messages. For accounts created on or after August 31, 2026 (00:00 UTC), replaying a block after an edit returns a 400 on the Claude API, Bedrock and Google Cloud:

messages.1.content.0: Invalid `signature` in `thinking` block. The block is bound to a different conversation. Remove the block, or set `thinking.block_binding.prefix_mismatch_behavior` to "drop_block".

Older accounts don’t enforce this by default, so a clean run on an old key proves nothing. Keep conversations append-only and change instructions or tools with mid-conversation system messages. If you must edit, send anthropic-beta: thinking-binding-controls-2026-08-01 and drop mismatched blocks:

"thinking": {"type": "adaptive",
  "block_binding": {"prefix_mismatch_behavior": "drop_block"}}

That works only with adaptive thinking; with between_tools, strip thinking blocks from the edited turn on. Sonnet 5.5 reads blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier; it drops Opus 5, Opus 5.5, Fable and Mythos blocks, and Sonnet 5.5 blocks from another account, without failing the request. No other model reads its blocks. See Claude Fable 5.1 preserved thinking.

4. computer_20251124 fails on the Claude API and Google Cloud

There, computer use needs the new toolset. The error begins:

'claude-sonnet-5-5' does not support tool types: computer_20251124.
// Before (claude-sonnet-5)
"tools": [{"type": "computer_20251124", ...}]

// After (claude-sonnet-5-5, Claude API and Google Cloud)
"tools": [{"type": "computer_toolset_20260801"}]

Drop the old computer-use beta header and update your loop for member tool_use blocks, batched actions and toolset_name on results. Bedrock still accepts computer_20251124; computer_20250124 fails everywhere.

5. Some advisor pairings are rejected

With the advisor tool (beta), a Sonnet 5.5 executor accepts only Opus 5, Opus 5.5, Sonnet 5.5, Fable 5, Fable 5.1, Mythos 5 or Mythos 5.1 as advisor. Sonnet 5, Opus 4.8 and Opus 4.7 advisors now return a 400. Advice also arrives encrypted as an advisor_redacted_result block, so code that parses advice text gets nothing.

The silent change: text between tool calls moves into thinking blocks

This one fails nothing. On Sonnet 5, notes between tool calls came back as text. On Sonnet 5.5, anything longer than a sentence or two arrives as a progress-update thinking block, empty under the default display: "omitted". An agent UI that streams those notes goes quiet, with no error. Three fixes:

// Header: anthropic-beta: thinking-display-updates-2026-08-18
"thinking": {"type": "adaptive", "display": "updates"}

Behavior changes with no code change

These don’t fail requests, but they change output and cost. The prompting guide has fixes.

Cost per task: same price, fewer dollars per result

Anthropic’s launch post says Sonnet 5.5 “costs up to 30% less per task than its predecessor” in its testing. Prices are identical, so the saving comes from fewer tokens and steps. Its per-effort charts show lower effort on 5.5 beating Sonnet 5’s best run:

Benchmark (Anthropic’s charts) Sonnet 5.5 Sonnet 5, best run
Terminal-Bench 4.0 28.8% at medium, $0.83 10.3% at max, $11.62
FrontierCode 1.1 49.4% at high, $0.42 42.7% at xhigh, $10.07
CursorBench 4.0 35.8% at low, $0.50 34.1% at max, $7.17

CursorBench costs are Anthropic’s estimates at list prices. The flip side is max: per OfficeChai’s write-up of Artificial Analysis data, Sonnet 5.5 at max uses about 193,000 output tokens per index task, and cost per task runs about 50% above Sonnet 5. The savings live at high and below; see Claude Sonnet 5.5 pricing.

Migration checklist

Swap claude-sonnet-5 for claude-sonnet-5-5, then run the six checks from Anthropic’s migration guide:

  1. Replace disabled with between_tools, at high effort or below.
  2. Replace forced tool_choice with auto, strict: true and a prompt line (on Bedrock, validate in code).
  3. Keep history append-only; use mid-conversation system messages for changes.
  4. Move computer use to computer_toolset_20260801 on the Claude API and Google Cloud.
  5. Pick a supported advisor, and stop parsing advice text.
  6. Set thinking.display if your UI shows text between tool calls.

Then re-run your effort sweep. Claude Code can automate the migration:

/claude-api migrate this project to claude-sonnet-5-5

The Opus equivalent is Claude Opus 5.5 vs Opus 5 migration.

Turn the migration into a regression test in Apidog

In Apidog, save three requests to https://api.anthropic.com/v1/messages in one project, all on the same environment:

  1. Baseline: your current Sonnet 5 body.
  2. Old body, new model: only model changed to claude-sonnet-5-5. Assert status 400 and an error message mentioning between_tools or tool_choice.
  3. Migrated: the fixed body. Assert status 200, a stop_reason that isn’t refusal or max_tokens, and a tool_use block for tool requests.

Keep ANTHROPIC_API_KEY in the environment and reference it as {{ANTHROPIC_API_KEY}} in the x-api-key header. Compare usage.output_tokens across effort levels for your own cost-per-task number. Saved as a test scenario, a regression to disabled fails the run instead of production. Request basics: how to use the Claude Sonnet 5.5 API.

FAQ

Is Claude Sonnet 5.5 more expensive than Sonnet 5? No. Both cost $2/$10 per million tokens with $0.20 cache reads, and the tokenizer is the same.

Why do I get “thinking.type.disabled is not supported” on Sonnet 5.5? disabled was removed. Send thinking: {"type": "between_tools"} at low, medium or high effort, with no other thinking fields.

Do Sonnet 5 conversations carry over to Sonnet 5.5? Yes. Sonnet 5.5 reads Sonnet 5 thinking blocks. Moving back loses them: no other model reads Sonnet 5.5 blocks.

Should I move everything to Sonnet 5.5? For most workloads, yes. Watch cyber-adjacent work, which may fall back to Sonnet 5, and Bedrock tool calls, which lose strict mode. Our Claude Sonnet 5 guide covers the model you’re leaving.

Your next step

Save your Sonnet 5 request and its 5.5 twin side by side, confirm the 400, fix it, and move traffic once the fix passes. Download Apidog to build that test.

Explore more

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

Is Claude Sonnet 5.5 free? Yes on Claude.ai (web, iOS, Android). Every free route checked, plus what isn't: Claude Code, the API, and Copilot Free.

29 September 2026

How to Use Claude Sonnet 5.5 in Claude Code (and When to Keep Opus 5.5)

How to Use Claude Sonnet 5.5 in Claude Code (and When to Keep Opus 5.5)

Claude Sonnet 5.5 Claude Code setup: v2.1.284+, claude --model claude-sonnet-5-5, effort levels, the sonnet alias trap, and when to keep Opus 5.5.

29 September 2026

Claude Sonnet 5.5 Pricing: The Full Cost Breakdown (API, Caching, Batch, and Plans)

Claude Sonnet 5.5 Pricing: The Full Cost Breakdown (API, Caching, Batch, and Plans)

Claude Sonnet 5.5 pricing: $2/$10 per million tokens, $0.20 cache reads, batch at half price. Worked cost examples, effort costs, and plan prices.

29 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Claude Sonnet 5.5 vs Sonnet 5: What Changed, and the Breaking Changes to Fix Before You Switch