Claude Sonnet 5.5 costs $2/$10 per million input/output tokens, half of Claude Opus 5.5’s $4/$20. On Anthropic’s launch table it lands within a few points of Opus 5.5 on most rows and beats it on Terminal-Bench 4.0 (70.6% vs 66.4%), yet Anthropic says Opus 5.5 “remains clearly stronger at complex, open-ended work.” The verdict: route well-scoped, high-volume work and subagents to Sonnet 5.5 at low to high effort. Pay for Opus 5.5 on long, open-ended tasks, or wherever you’d push Sonnet to xhigh or max, because Opus at medium or high often scores higher for similar or less money.
For each model alone, see what is Claude Sonnet 5.5 and what is Claude Opus 5.5. The last section shows how to test both on your own prompts in Apidog.
Specs and price side by side
| Sonnet 5.5 | Opus 5.5 | |
|---|---|---|
| API model ID | claude-sonnet-5-5 |
claude-opus-5-5 |
| Input / output, per 1M tokens | $2 / $10 | $4 / $20 |
| Cache write (5 min) | $2.50 | $5 |
| Cache read | $0.20 | $0.20 |
| Fast mode | Not available | $8 / $40 |
| Default effort on the API | high |
medium |
| Thinking | Adaptive by default, or between_tools |
Adaptive, always on |
| Context / max output | 1M / 128K | 1M / 128K |
| Latency class | Fast | Moderate |
Three rows matter most:
- Cache reads cost the same, so in cache-heavy agent loops the gap sits mostly in output and cache writes. Claude Sonnet 5.5 pricing does the math. Neither model charges a long-context premium.
- Default effort differs. A default-settings test pits Sonnet at
highagainst Opus atmedium, the pairing where Opus tends to win. - Fast mode is Opus-only (docs). Anthropic says Sonnet 5.5 generates output 30%+ faster than Sonnet 5.
Benchmarks: close on the launch table, wider on the system card
Rows one to eight are Anthropic’s launch table; the rest come from the system card. Cognition ran FrontierCode, Cursor ran CursorBench, Zapier ran AutomationBench, and Artificial Analysis (AA) ran GDPval-AA and AA-Briefcase; Anthropic ran the others. Sonnet is at max effort unless noted.
| Benchmark | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% (xhigh) |
| FrontierCode 1.1 Main | 46.2% max, 52.1% xhigh | 54.4% |
| CursorBench 4.0 | 55.5% | 57.8% |
| GDPval-AA v2.1 (Elo) | 1844 | 1846 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1822 |
| HLE, with tools | 64.5% | 67.7% |
| OSWorld 2.1, partial | 80.1% | 81.8% |
| Chartography, no tools | 61.6% | 64.4% |
| SWE-Bench Pro | 81.3 | 89.9 |
| SWE-Bench Multimodal | 54.3 | 61.4 |
| HLE, no tools | 56.9 | 64.4 |
| OSWorld 2.1, strict pass | 43.5% | 48.7% |
| ProgramBench, long context | 79.7% | 91.2% |
| Terminal-Bench-Science 0.1 | 59.9% | 58.7% |
| AutomationBench | 44.7 | 42.5 |
| HealthBench Professional | 69.2 | 65.6 |
The launch table is the flattering cut: outside Terminal-Bench, Opus leads its percentage rows by 1.7 to 3.2 points, counting Sonnet’s xhigh FrontierCode score. Five system card rows widen the gap to 5.2 to 11.5 points: SWE-Bench Pro, SWE-Bench Multimodal, HLE without tools, strict OSWorld and long-context ProgramBench. Long, unassisted, whole-task work is where Opus stays ahead.
Read the Terminal-Bench 4.0 win with care: system card section 8.5 says safeguard fallback touched 10% of Opus trials against 1.5% of Sonnet’s, and AA’s own run scored Sonnet 5.5 at 63.6%. Details in Claude Sonnet 5.5 benchmarks.
Effort decides the matchup
The launch table compares each model at its best setting. Your bill doesn’t. Anthropic’s launch page plots every effort level with cost per attempt or task:
| Benchmark, model | Low | Medium | High | Xhigh | Max |
|---|---|---|---|---|---|
| Terminal-Bench, Sonnet | 20.0% ($0.76) | 28.8% ($0.83) | 43.0% ($1.94) | 61.5% ($5.30) | 70.6% ($12.54) |
| Terminal-Bench, Opus | 38.5% ($1.29) | 57.6% ($2.94) | 64.2% ($3.88) | 66.4% ($7.35) | 64.8% ($11.24) |
| CursorBench, Sonnet | 35.8% ($0.50) | 39.2% ($0.70) | 47.8% ($1.67) | 53.1% ($3.88) | 55.5% ($9.67) |
| CursorBench, Opus | 43.7% ($1.17) | 52.5% ($2.91) | 56.0% ($3.97) | 56.0% ($6.98) | 57.8% ($13.43) |
| FrontierCode, Sonnet | 29.3% ($0.19) | 36.5% ($0.24) | 49.4% ($0.42) | 52.1% ($1.59) | 46.2% ($20.78) |
| FrontierCode, Opus | 47.3% ($0.40) | 54.6% ($0.80) | 54.0% ($1.09) | 51.4% ($2.25) | 54.4% ($6.19) |
| AA-Briefcase, Sonnet | 1264 ($0.87) | 1461 ($1.64) | 1634 ($3.95) | 1746 ($9.63) | 1811 ($29.19) |
| AA-Briefcase, Opus | 1285 ($1.15) | 1642 ($4.40) | 1705 ($6.27) | 1780 ($12.27) | 1822 ($21.05) |
What it shows:
- Opus at medium often beats Sonnet at high. Terminal-Bench 4.0: Opus scores 57.6% for $2.94 per attempt, Sonnet 43.0% for $1.94.
- Opus below max can beat Sonnet’s best for less. CursorBench: Opus at high (56.0%, $3.97) vs Sonnet at max (55.5%, $9.67). FrontierCode: Opus at medium (54.6%, $0.80) vs Sonnet’s best (52.1%, $1.59).
- Sonnet at max can cost more than Opus at max: $29.19 vs $21.05 on AA-Briefcase, for a lower score. On FrontierCode, Sonnet’s max score fell below its xhigh score because it fanned out review subagents, Anthropic’s footnote says.
- Sonnet owns the bottom of the curve. Its low setting undercuts Opus’s cheapest point on every chart.
Half the token price isn’t half the cost per task: Sonnet spends more tokens as effort climbs. AA data, as OfficeChai reported it, puts Sonnet 5.5 at about 193,000 output tokens per index task at max, roughly 60% more than Opus 5.5.
That’s the main debate in the Hacker News thread: many commenters read the charts as saying Opus at lower effort matches or undercuts Sonnet at high or xhigh, leaving Sonnet’s niche at low or medium effort and as a subagent under an Opus orchestrator.
Anthropic’s prompting guide says Sonnet 5.5’s effort is recalibrated: start at high (medium for well-specified agentic coding) and save xhigh and max for measured gains. Changing top-level effort between requests invalidates the prompt cache, so set it per workload (API guide).
Safety and behavior differences
Prompt injection results split. On Gray Swan’s indirect prompt injection benchmark, Sonnet 5.5’s attack success rate is 3.4% at 15 attempts (Sonnet 5: 6.7%): less resistant than Opus 5.5, more than every non-Claude model tested, weakest in GUI computer use (12.5%). Against Shade’s adaptive attackers, Sonnet beats Opus in coding (3.01% vs 54.61% without safeguards, 2.63% vs 11.13% with probes on) and ties it in computer use (0.07%); in browser use it’s the first model Anthropic evaluated with no successful attacks. See Opus 5.5 and prompt injection.
Both models carry cyber safeguards; Sonnet 5.5 is the first Sonnet to get them. Flagged higher-risk cyber tasks fall back to Sonnet 5, automatically in Anthropic’s apps and on the API only if you opt in (fallbacks: "default", beta). The system card warns of more refusals even on benign security work.
The system card also finds Sonnet 5.5 more honest under pressure than Opus 5.5 but more prone to hallucination.
Mixing Sonnet 5.5 and Opus 5.5 in one system
One way to get both: Opus plans, Sonnet executes. Three mechanics matter.
Thinking blocks don’t cross. Sonnet 5.5 drops Opus 5 and Opus 5.5 thinking blocks (unbilled; the request still returns 200), and blocks are bound to model, conversation and account. Switch models mid-conversation and the reasoning is gone, so hand off through text: Opus writes the plan as a message, Sonnet starts from it (migration guide).
The advisor tool accepts Opus 5.5 as advisor to a Sonnet 5.5 executor; advice returns encrypted as advisor_redacted_result.
Claude Code has opusplan: Opus in plan mode, Sonnet for execution. The sonnet alias means Sonnet 5.5 only on the Anthropic API (v2.1.284 or later), so on Bedrock, Google Cloud and Foundry opusplan executes on an older Sonnet. See Claude Sonnet 5.5 in Claude Code.
Where GPT-6 Sol fits
GPT-6 Sol shares Sonnet 5.5’s $2/$10 list price, with $0.20 cached reads (90% off) and an 872k context window. Anthropic’s table gives Sol four cells: FrontierCode 49.3%, GDPval-AA 1487, AA-Briefcase 1483 and Chartography 53.6% without tools. A footnote says OpenAI recently fixed a Sol image-understanding bug that the AA and Surge AI scores may not reflect yet.
Cost per task flips by benchmark. On AA-Briefcase at max, Sol costs $2.67 per task against Sonnet’s $29.19, scoring 1483 to 1811. On FrontierCode, Sonnet at high matches Sol’s best (49.4% vs 49.3%) for $0.42 against $2.07. See GPT-6 Sol vs Claude Opus 5.5 and what is GPT-6 Sol.
Which model for which workload
| Workload | Model and effort |
|---|---|
| High-volume chat, classification, extraction | Sonnet 5.5, low or medium |
| Well-scoped bug fixes, PR-sized changes | Sonnet 5.5, medium or high |
| Subagents running an Opus plan | Sonnet 5.5, medium or high |
| Long, open-ended agentic work | Opus 5.5, medium then high |
Anything needing Sonnet at xhigh or max |
Opus 5.5, medium or high |
| Long-context codebase reasoning | Opus 5.5, medium or above |
| Agents reading untrusted content | Either; test your own injection cases |
Test both on your own traffic
Every chart above came from someone else’s harness, not your prompts, tools or token mix. Start with the pairing the data flags:
ask() {
curl -s https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"'"$1"'","max_tokens":16000,
"output_config":{"effort":"'"$2"'"},
"messages":[{"role":"user","content":"Review this diff and list each bug as JSON: ..."}]}' \
| jq '{model, stop_reason, usage}'
}
ask claude-sonnet-5-5 high
ask claude-opus-5-5 medium
In Apidog, make it repeatable: one request to https://api.anthropic.com/v1/messages, ANTHROPIC_API_KEY as an environment variable, and {{model}} and {{effort}} in the body. Duplicate it per pairing and add a post-processor so every run checks the same things:
const body = pm.response.json();
pm.test("finished cleanly", () => {
pm.expect(body.stop_reason).to.not.be.oneOf(["max_tokens", "refusal"]);
});
pm.test("stays inside the token budget", () => {
pm.expect(body.usage.output_tokens).to.be.below(8000);
});
Run each pair 10 to 20 times and compare usage, response time and pass rate; tokens times list price is your real cost per task. More in how to test LLM applications.
FAQ
Is Claude Sonnet 5.5 better than Opus 5.5? Not on most rows. Opus 5.5 leads the launch table except on Terminal-Bench 4.0 and a GDPval-AA near-tie. Last generation’s matchup: Claude Opus 5 vs Sonnet 5.
Is Sonnet 5.5 half the cost of Opus 5.5? Per token, yes. Per task it depends on effort: at max, Sonnet cost more on AA-Briefcase ($29.19 vs $21.05).
Can I switch models mid-conversation? Yes, but Sonnet 5.5 drops Opus 5.5’s thinking blocks, so pass state as text.
Sonnet 5.5 or GPT-6 Sol at $2/$10? Sonnet leads all four shared rows at its best setting; Sol can be far cheaper per task. Test both.
Next step
Run your two most expensive prompts through Sonnet 5.5 at high and Opus 5.5 at medium, and compare pass rate and cost per task before you change a routing rule. To keep the requests, assertions and results in one project, download Apidog.



