Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, 55.5% on CursorBench 4.0, 46.2% on FrontierCode 1.1 at max effort (52.1% at xhigh) and 81.3 on SWE-Bench Pro, up from Sonnet 5’s 10.3%, 34.1%, 42.4% and 63.2. Opus 5.5 leads most launch-table rows by about 2 to 3 points but trails on Terminal-Bench. Artificial Analysis independently scores Sonnet 5.5 at 56 on its Intelligence Index, #3 of 216.
Below: the Claude Sonnet 5.5 benchmarks in full, who ran them, what effort changes, and how to test the model on your own traffic in Apidog. For specs, see what is Claude Sonnet 5.5.
The launch table
Anthropic’s launch post compares four models (“n/r” marks a dash). Sonnet 5.5 cells are max effort unless marked; the API defaults to high, Claude Code and the apps to medium.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4%¹ | n/r |
| FrontierCode 1.1 (Main) | 46.2% Max² / 52.1% Xhigh | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | n/r |
| GDPval-AA v2.1³ | 1844 | 1449 | 1846 | 1487⁴ |
| AA-Briefcase v1.1³ | 1811 | 1359 | 1822 | 1483⁴ |
| Humanity’s Last Exam (with tools) | 64.5% | 54.9% | 67.7% | n/r |
| OSWorld 2.1 (partial) | 80.1% | 57.0% | 81.8% | n/r |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6%⁴ |
The footnotes, in plain terms:
- Opus 5.5’s Terminal-Bench score is at xhigh, its best; at max it scored 64.8%.
- FrontierCode penalizes out-of-scope edits. At max, Sonnet 5.5 more often ran Claude Code’s code-review skill, fanning out to many subagents; in two cases Cognition examined, that caused a timeout or out-of-scope edits.
- Artificial Analysis ran both knowledge-work tests on a pre-release deployment with a structured-output bug, since fixed. Anthropic expects a small effect, if anything understating Sonnet 5.5.
- OpenAI recently fixed a GPT-6 Sol image-understanding bug that the AA and Surge AI scores may not reflect.
Who ran each benchmark
Anthropic ran every test itself unless a third party is named. Cross-vendor rows share a benchmark name, not a harness or effort setting.
| Runner | Benchmark | Conditions |
|---|---|---|
| Anthropic | Terminal-Bench 4.0, HLE, OSWorld 2.1 | TB: Claude Code --bare, 5 trials, safeguards on. HLE: search and code execution. OSWorld: 108 tasks |
| Cognition | FrontierCode 1.1 | Claude in Claude Code, GPT in Codex CLI |
| Cursor | CursorBench 4.0 | Cursor’s production harness; Anthropic estimated cost at list price |
| Artificial Analysis | GDPval-AA, AA-Briefcase | Pre-release deployment |
| Anthropic | Chartography | Surge AI’s benchmark, graded with Gemini 3.5 Flash; Sol’s score from Surge |
Fallback touched 1.5% of Sonnet 5.5’s Terminal-Bench trials but 10% of Opus 5.5’s (system card §8.5), so read Sonnet’s 4.2-point lead with care.
System card extras
Rows one to six are the system card’s max-effort summary.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| SWE-Bench Pro | 81.3 | 63.2 | 89.9 |
| SWE-Bench Multilingual | 90.3 | 78.3 | 93.9 |
| SWE-Bench Multimodal | 54.3 | 28.1 | 61.4 |
| HLE, no tools | 56.9 | 43.1 | 64.4 |
| HealthBench Professional | 69.2 | 57.8 | 65.6 |
| AutomationBench (Zapier’s run) | 44.7 | 10.7 | 42.5 |
| DeepSWE v1.1 | 71.0% | n/r | n/r |
| Terminal-Bench-Science 0.1 | 59.9% | n/r | 58.7% |
| FrontierSWE v2 (Proximal’s run) | 61.9% | n/r | 62.3% |
| ArXivMath (no tools / tools) | 86.8% / 95.2% | n/r | 91.2% / 96.9% |
| ProgramBench (long context) | 79.7% | 77.3% | 91.2% |
| OSWorld 2.1, strict pass | 43.5% | 25.6% | 48.7% |
| Toolathlon-Verified (Pass@1) | 77.8% | 74.7% | 77.8% |
| OfficeQA / OfficeQA Pro | 76.9% / 65.6% | 75.1% / 62.1% | 78.9% / 67.7% |
Sonnet 5.5 edges Opus 5.5 on HealthBench Professional, AutomationBench and Terminal-Bench-Science; Opus leads widest on ProgramBench, SWE-Bench Pro and HLE without tools. OSWorld’s 80.1% headline is partial credit; only 43.5% of tasks pass outright.
Effort changes the picture
The launch charts, as score (cost). Terminal-Bench cost is per attempt, the others per task. Two charts show GPT-5.6 Sol (*) because GPT-6 Sol numbers weren’t published.
| Model | Low | Medium | High | Xhigh | Max |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | |||||
| Sonnet 5.5 | 20.0% ($0.76) | 28.8% ($0.83) | 43.0% ($1.94) | 61.5% ($5.30) | 70.6% ($12.54) |
| Opus 5.5 | 38.5% ($1.29) | 57.6% ($2.94) | 64.2% ($3.88) | 66.4% ($7.35) | 64.8% ($11.24) |
| Sonnet 5 | 3.2% ($3.78) | 4.4% ($4.83) | 4.5% ($8.20) | 7.0% ($9.95) | 10.3% ($11.62) |
| GPT-5.6 Sol* | 7.9% ($1.46) | 20.9% ($2.69) | 26.1% ($4.12) | 28.5% ($5.39) | 37.3% ($7.89) |
| FrontierCode 1.1 Main | |||||
| Sonnet 5.5 | 29.3% ($0.19) | 36.5% ($0.24) | 49.4% ($0.42) | 52.1% ($1.59) | 46.2% ($20.78) |
| Opus 5.5 | 47.3% ($0.40) | 54.6% ($0.80) | 54.0% ($1.09) | 51.4% ($2.25) | 54.4% ($6.19) |
| Sonnet 5 | 28.7% ($2.39) | 35.2% ($3.81) | 39.4% ($6.10) | 42.7% ($10.07) | 42.4% ($17.12) |
| GPT-6 Sol | 37.3% ($0.43) | 45.9% ($0.77) | 47.7% ($1.04) | 48.4% ($1.32) | 49.3% ($2.07) |
| CursorBench 4.0 | |||||
| Sonnet 5.5 | 35.8% ($0.50) | 39.2% ($0.70) | 47.8% ($1.67) | 53.1% ($3.88) | 55.5% ($9.67) |
| Opus 5.5 | 43.7% ($1.17) | 52.5% ($2.91) | 56.0% ($3.97) | 56.0% ($6.98) | 57.8% ($13.43) |
| Sonnet 5 | 24.1% ($1.39) | 28.0% ($2.31) | 30.8% ($3.48) | 32.0% ($4.55) | 34.1% ($7.17) |
| GPT-5.6 Sol* | 24.6% ($0.87) | 31.1% ($1.77) | 35.7% ($2.85) | 37.7% ($4.40) | 41.7% ($8.23) |
| AA-Briefcase v1.1 (Elo) | |||||
| Sonnet 5.5 | 1264 ($0.87) | 1461 ($1.64) | 1634 ($3.95) | 1746 ($9.63) | 1811 ($29.19) |
| Opus 5.5 | 1285 ($1.15) | 1642 ($4.40) | 1705 ($6.27) | 1780 ($12.27) | 1822 ($21.05) |
| Sonnet 5 | 923 ($0.82) | 1056 ($1.73) | 1177 ($3.76) | 1274 ($7.56) | 1359 ($14.43) |
| GPT-6 Sol | 905 ($0.12) | 1142 ($0.34) | 1289 ($0.63) | 1364 ($1.19) | 1483 ($2.67) |
- Most gains arrive by xhigh. Max adds 9.1 points on Terminal-Bench for 2.4x the cost, 2.4 on CursorBench for 2.5x.
- FrontierCode max scores lower and costs 13x more than xhigh (46.2% at $20.78 vs 52.1% at $1.59).
- Opus 5.5 at lower effort is the real rival. On Terminal-Bench, Opus at medium gets 57.6% for $2.94; Sonnet at high gets 43.0% for $1.94. Opus at high (64.2%, $3.88) beats Sonnet at xhigh (61.5%, $5.30). See Sonnet 5.5 vs Opus 5.5.
- High is Sonnet’s cost sweet spot. FrontierCode at high: 49.4% for $0.42, matching GPT-6 Sol’s best (49.3%) at about a fifth of the cost (pricing breakdown).
Independent results
Artificial Analysis scores Sonnet 5.5 at 56 on Intelligence Index v4.3.2, #3 of 216, behind only Opus 5.5 at max and xhigh. Sonnet 5 scored 38. Cost to run the index:
| Effort | Index score | Cost to run |
|---|---|---|
| max | 55.98 | $8,977 |
| xhigh | 51.85 | $2,738 |
| high | 46.74 | $1,176 |
| medium | 40.74 | $701 |
| low | 35.84 | $544 |
Max costs 7.6x high for 9.2 more points. OfficeChai’s reading of AA puts Sonnet 5.5 off the Pareto frontier, most cost-competitive at high, and counts about 193,000 output tokens per index task at max.
- AA’s own Terminal-Bench 4.0 run: 63.6%, versus Anthropic’s 70.6%.
- Sonnet 5’s low score is real: AA measured 14.1% (see Sonnet 5 benchmarks).
- AA-Omniscience (per OfficeChai): 54% accuracy and a 47% hallucination rate, versus 66% and 59% for Opus 5.5.
- Not measured yet: output speed and time to first token, so “30%+ faster” remains Anthropic’s claim. No LMArena listing yet.
Safety benchmarks
Prompt injection, per the system card:
- Gray Swan: attack success 0.4%, 2.7% and 3.4% at k=1, 10 and 15 (Sonnet 5: 0.7%, 5.1%, 6.7%). GUI computer use is weakest at 12.5% (k=15).
- Shade, coding: 3.01% without safeguards, mostly via the Sonnet 5 cyber fallback; Sonnet 5.5 itself was compromised on 4 of 5,901 requests.
- Browser use: no attack succeeded across 110 scenarios.
Cyber evals ran with safeguards off: 46.1% on CyScenarioBench (Sonnet 5: 0.7%, Opus 5.5: 67.6%). It’s the first Sonnet with cyber safeguards, and Anthropic warns of more refusals even on benign security tasks. See our prompt injection guide.
Run your own eval
Benchmarks don’t look like your traffic, and Anthropic’s prompting guide says effort is recalibrated, so don’t reuse Sonnet 5 settings. In Apidog:
- Store
ANTHROPIC_API_KEYas an environment variable and save 20 to 50 real prompts as Messages requests. - Assert status 200, a
stop_reasonother than"max_tokens"or"refusal", and the content a correct answer needs. - Run it at
low,medium,highandxhigh, one pass each; changing effort invalidates the prompt cache. - Record pass rate and the token counts in
usage, priced at $2 input and $10 output per million.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"max_tokens": 16000,
"output_config": {"effort": "high"},
"messages": [{"role": "user", "content": "Classify this support ticket: ..."}]
}'
Ship the lowest effort whose pass rate you accept. See testing LLM applications and testing AI agent APIs, then download Apidog to build the collection.
FAQ
Does Sonnet 5.5 beat Opus 5.5? On Terminal-Bench 4.0, HealthBench Professional, AutomationBench, Terminal-Bench-Science and Chartography with tools, yes. Opus 5.5 leads or ties the rest.
What is Sonnet 5.5’s SWE-Bench score? 81.3 on SWE-Bench Pro and 90.3 on Multilingual, at max effort.
Why are there two FrontierCode numbers? 46.2% is max, 52.1% is xhigh; max’s subagent fan-out drew out-of-scope penalties.
Should I upgrade from Sonnet 5? Yes on benchmarks, but read the breaking API changes in Sonnet 5.5 vs Sonnet 5 first.
Next step
Start at high (or medium for well-specified agentic coding), as Anthropic suggests, and let your eval in Apidog decide whether a step up pays.



