Claude Sonnet 5.5 Benchmarks: Anthropic's Numbers, Independent Results, and What They Mean

Claude Sonnet 5.5 benchmarks: 70.6% Terminal-Bench 4.0, 81.3 SWE-Bench Pro, AA index 56. Who ran each test, what effort costs, how to run your own eval.

Medy Evrard

29 September 2026

Claude Sonnet 5.5 Benchmarks: Anthropic's Numbers, Independent Results, and What They Mean

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, 55.5% on CursorBench 4.0, 46.2% on FrontierCode 1.1 at max effort (52.1% at xhigh) and 81.3 on SWE-Bench Pro, up from Sonnet 5’s 10.3%, 34.1%, 42.4% and 63.2. Opus 5.5 leads most launch-table rows by about 2 to 3 points but trails on Terminal-Bench. Artificial Analysis independently scores Sonnet 5.5 at 56 on its Intelligence Index, #3 of 216.

Below: the Claude Sonnet 5.5 benchmarks in full, who ran them, what effort changes, and how to test the model on your own traffic in Apidog. For specs, see what is Claude Sonnet 5.5.

The launch table

Anthropic’s launch post compares four models (“n/r” marks a dash). Sonnet 5.5 cells are max effort unless marked; the API defaults to high, Claude Code and the apps to medium.

Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0 70.6% 10.3% 66.4%¹ n/r
FrontierCode 1.1 (Main) 46.2% Max² / 52.1% Xhigh 42.4% 54.4% 49.3%
CursorBench 4.0 55.5% 34.1% 57.8% n/r
GDPval-AA v2.1³ 1844 1449 1846 1487⁴
AA-Briefcase v1.1³ 1811 1359 1822 1483⁴
Humanity’s Last Exam (with tools) 64.5% 54.9% 67.7% n/r
OSWorld 2.1 (partial) 80.1% 57.0% 81.8% n/r
Chartography (no tools) 61.6% 15.6% 64.4% 53.6%⁴

The footnotes, in plain terms:

  1. Opus 5.5’s Terminal-Bench score is at xhigh, its best; at max it scored 64.8%.
  2. FrontierCode penalizes out-of-scope edits. At max, Sonnet 5.5 more often ran Claude Code’s code-review skill, fanning out to many subagents; in two cases Cognition examined, that caused a timeout or out-of-scope edits.
  3. Artificial Analysis ran both knowledge-work tests on a pre-release deployment with a structured-output bug, since fixed. Anthropic expects a small effect, if anything understating Sonnet 5.5.
  4. OpenAI recently fixed a GPT-6 Sol image-understanding bug that the AA and Surge AI scores may not reflect.

Who ran each benchmark

Anthropic ran every test itself unless a third party is named. Cross-vendor rows share a benchmark name, not a harness or effort setting.

Runner Benchmark Conditions
Anthropic Terminal-Bench 4.0, HLE, OSWorld 2.1 TB: Claude Code --bare, 5 trials, safeguards on. HLE: search and code execution. OSWorld: 108 tasks
Cognition FrontierCode 1.1 Claude in Claude Code, GPT in Codex CLI
Cursor CursorBench 4.0 Cursor’s production harness; Anthropic estimated cost at list price
Artificial Analysis GDPval-AA, AA-Briefcase Pre-release deployment
Anthropic Chartography Surge AI’s benchmark, graded with Gemini 3.5 Flash; Sol’s score from Surge

Fallback touched 1.5% of Sonnet 5.5’s Terminal-Bench trials but 10% of Opus 5.5’s (system card §8.5), so read Sonnet’s 4.2-point lead with care.

System card extras

Rows one to six are the system card’s max-effort summary.

Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5
SWE-Bench Pro 81.3 63.2 89.9
SWE-Bench Multilingual 90.3 78.3 93.9
SWE-Bench Multimodal 54.3 28.1 61.4
HLE, no tools 56.9 43.1 64.4
HealthBench Professional 69.2 57.8 65.6
AutomationBench (Zapier’s run) 44.7 10.7 42.5
DeepSWE v1.1 71.0% n/r n/r
Terminal-Bench-Science 0.1 59.9% n/r 58.7%
FrontierSWE v2 (Proximal’s run) 61.9% n/r 62.3%
ArXivMath (no tools / tools) 86.8% / 95.2% n/r 91.2% / 96.9%
ProgramBench (long context) 79.7% 77.3% 91.2%
OSWorld 2.1, strict pass 43.5% 25.6% 48.7%
Toolathlon-Verified (Pass@1) 77.8% 74.7% 77.8%
OfficeQA / OfficeQA Pro 76.9% / 65.6% 75.1% / 62.1% 78.9% / 67.7%

Sonnet 5.5 edges Opus 5.5 on HealthBench Professional, AutomationBench and Terminal-Bench-Science; Opus leads widest on ProgramBench, SWE-Bench Pro and HLE without tools. OSWorld’s 80.1% headline is partial credit; only 43.5% of tasks pass outright.

Effort changes the picture

The launch charts, as score (cost). Terminal-Bench cost is per attempt, the others per task. Two charts show GPT-5.6 Sol (*) because GPT-6 Sol numbers weren’t published.

Model Low Medium High Xhigh Max
Terminal-Bench 4.0
Sonnet 5.5 20.0% ($0.76) 28.8% ($0.83) 43.0% ($1.94) 61.5% ($5.30) 70.6% ($12.54)
Opus 5.5 38.5% ($1.29) 57.6% ($2.94) 64.2% ($3.88) 66.4% ($7.35) 64.8% ($11.24)
Sonnet 5 3.2% ($3.78) 4.4% ($4.83) 4.5% ($8.20) 7.0% ($9.95) 10.3% ($11.62)
GPT-5.6 Sol* 7.9% ($1.46) 20.9% ($2.69) 26.1% ($4.12) 28.5% ($5.39) 37.3% ($7.89)
FrontierCode 1.1 Main
Sonnet 5.5 29.3% ($0.19) 36.5% ($0.24) 49.4% ($0.42) 52.1% ($1.59) 46.2% ($20.78)
Opus 5.5 47.3% ($0.40) 54.6% ($0.80) 54.0% ($1.09) 51.4% ($2.25) 54.4% ($6.19)
Sonnet 5 28.7% ($2.39) 35.2% ($3.81) 39.4% ($6.10) 42.7% ($10.07) 42.4% ($17.12)
GPT-6 Sol 37.3% ($0.43) 45.9% ($0.77) 47.7% ($1.04) 48.4% ($1.32) 49.3% ($2.07)
CursorBench 4.0
Sonnet 5.5 35.8% ($0.50) 39.2% ($0.70) 47.8% ($1.67) 53.1% ($3.88) 55.5% ($9.67)
Opus 5.5 43.7% ($1.17) 52.5% ($2.91) 56.0% ($3.97) 56.0% ($6.98) 57.8% ($13.43)
Sonnet 5 24.1% ($1.39) 28.0% ($2.31) 30.8% ($3.48) 32.0% ($4.55) 34.1% ($7.17)
GPT-5.6 Sol* 24.6% ($0.87) 31.1% ($1.77) 35.7% ($2.85) 37.7% ($4.40) 41.7% ($8.23)
AA-Briefcase v1.1 (Elo)
Sonnet 5.5 1264 ($0.87) 1461 ($1.64) 1634 ($3.95) 1746 ($9.63) 1811 ($29.19)
Opus 5.5 1285 ($1.15) 1642 ($4.40) 1705 ($6.27) 1780 ($12.27) 1822 ($21.05)
Sonnet 5 923 ($0.82) 1056 ($1.73) 1177 ($3.76) 1274 ($7.56) 1359 ($14.43)
GPT-6 Sol 905 ($0.12) 1142 ($0.34) 1289 ($0.63) 1364 ($1.19) 1483 ($2.67)

Independent results

Artificial Analysis scores Sonnet 5.5 at 56 on Intelligence Index v4.3.2, #3 of 216, behind only Opus 5.5 at max and xhigh. Sonnet 5 scored 38. Cost to run the index:

Effort Index score Cost to run
max 55.98 $8,977
xhigh 51.85 $2,738
high 46.74 $1,176
medium 40.74 $701
low 35.84 $544

Max costs 7.6x high for 9.2 more points. OfficeChai’s reading of AA puts Sonnet 5.5 off the Pareto frontier, most cost-competitive at high, and counts about 193,000 output tokens per index task at max.

Safety benchmarks

Prompt injection, per the system card:

Cyber evals ran with safeguards off: 46.1% on CyScenarioBench (Sonnet 5: 0.7%, Opus 5.5: 67.6%). It’s the first Sonnet with cyber safeguards, and Anthropic warns of more refusals even on benign security tasks. See our prompt injection guide.

Run your own eval

Benchmarks don’t look like your traffic, and Anthropic’s prompting guide says effort is recalibrated, so don’t reuse Sonnet 5 settings. In Apidog:

  1. Store ANTHROPIC_API_KEY as an environment variable and save 20 to 50 real prompts as Messages requests.
  2. Assert status 200, a stop_reason other than "max_tokens" or "refusal", and the content a correct answer needs.
  3. Run it at low, medium, high and xhigh, one pass each; changing effort invalidates the prompt cache.
  4. Record pass rate and the token counts in usage, priced at $2 input and $10 output per million.
curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "max_tokens": 16000,
    "output_config": {"effort": "high"},
    "messages": [{"role": "user", "content": "Classify this support ticket: ..."}]
  }'

Ship the lowest effort whose pass rate you accept. See testing LLM applications and testing AI agent APIs, then download Apidog to build the collection.

FAQ

Does Sonnet 5.5 beat Opus 5.5? On Terminal-Bench 4.0, HealthBench Professional, AutomationBench, Terminal-Bench-Science and Chartography with tools, yes. Opus 5.5 leads or ties the rest.

What is Sonnet 5.5’s SWE-Bench score? 81.3 on SWE-Bench Pro and 90.3 on Multilingual, at max effort.

Why are there two FrontierCode numbers? 46.2% is max, 52.1% is xhigh; max’s subagent fan-out drew out-of-scope penalties.

Should I upgrade from Sonnet 5? Yes on benchmarks, but read the breaking API changes in Sonnet 5.5 vs Sonnet 5 first.

Next step

Start at high (or medium for well-specified agentic coding), as Anthropic suggests, and let your eval in Apidog decide whether a step up pays.

Explore more

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

Is Claude Sonnet 5.5 free? Yes on Claude.ai (web, iOS, Android). Every free route checked, plus what isn't: Claude Code, the API, and Copilot Free.

29 September 2026

How to Use Claude Sonnet 5.5 in Claude Code (and When to Keep Opus 5.5)

How to Use Claude Sonnet 5.5 in Claude Code (and When to Keep Opus 5.5)

Claude Sonnet 5.5 Claude Code setup: v2.1.284+, claude --model claude-sonnet-5-5, effort levels, the sonnet alias trap, and when to keep Opus 5.5.

29 September 2026

Claude Sonnet 5.5 Pricing: The Full Cost Breakdown (API, Caching, Batch, and Plans)

Claude Sonnet 5.5 Pricing: The Full Cost Breakdown (API, Caching, Batch, and Plans)

Claude Sonnet 5.5 pricing: $2/$10 per million tokens, $0.20 cache reads, batch at half price. Worked cost examples, effort costs, and plan prices.

29 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Claude Sonnet 5.5 Benchmarks: Anthropic's Numbers, Independent Results, and What They Mean