Claude Haiku 5.5 scores 72.4% on OSWorld 2.1, 1620 on GDPval-AA v2.1, 46.4% on FrontierCode 1.1 and 39.2% on Terminal-Bench 4.0. Haiku 4.5 scored 15.7%, 735 and 0.0% on three of those. All are max-effort numbers; at the default medium effort, GDPval-AA drops to 1277. No independent lab has published Haiku 5.5 results yet.
Below: every Claude Haiku 5.5 benchmark, who ran it, how effort moves score and cost, and how to test the model on your own traffic in Apidog. For specs and pricing, start with what is Claude Haiku 5.5.
The launch table
Every Haiku 5.5 cell uses adaptive thinking at max effort, usually averaged over five trials (Terminal-Bench used ten per task). “n/r” means no result was published.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (Elo) | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam, no tools | 45.9% | 10.2% | n/r | 56.9% |
| Humanity’s Last Exam, with tools | 57.4% | 18.7% | n/r | 64.5% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | n/r | 42.4% | 52.1% (xhigh) |
| Chartography, no tools | 46.4% | 6.4% | 29.1% | 61.6% |
Who ran each benchmark
Anthropic ran most tests itself. Cross-vendor rows share a benchmark name, not always a harness or effort setting.
| Benchmark | Who ran it | Conditions |
|---|---|---|
| GDPval-AA v2.1, AA-Briefcase v1.1 | Artificial Analysis, independently | GDPval-AA: 220 tasks, 44 occupations, Elo anchored to DeepSeek V4.1 Flash (max) at 1600; Briefcase: multi-week projects |
| FrontierCode 1.1 | Cognition | Claude in Claude Code, GPT in Codex; Main = hardest 100 of 150 tasks |
| Chartography | Anthropic, on Surge AI’s benchmark | Gemini 3.5 Flash grader; Luna score from Surge AI |
| Terminal-Bench 4.0 | Anthropic | Claude Code --bare, 66 tasks, 10 trials each, safeguards on |
| OSWorld 2.1 | Anthropic | 82 of 108 tasks, 1080p, up to 500 steps |
| Humanity’s Last Exam | Anthropic | 980K-token task budget, Opus 4.6 grader |
Two GPT-6 Luna cells need context. Anthropic ran Luna’s OSWorld score through OpenAI’s API on the same 82 tasks. Luna’s Terminal-Bench 16.4% comes from the public leaderboard (Codex CLI, max effort). On Terminal-Bench, safeguards stopped 1.8% of Haiku 5.5 trials (12 of 660), and each counted as a fail.
The FrontierCode trap
The launch table puts Haiku 5.5 at max (46.4%) next to Sonnet 5.5 at xhigh (52.1%), Sonnet’s best score. The system card gives the like-for-like numbers:
| FrontierCode 1.1 | Main | Extended |
|---|---|---|
| Haiku 5.5, max | 46.4% | 58.4% |
| Haiku 5.5, xhigh | 45.8% | n/r |
| Sonnet 5.5, max | 46.2% | 59.1% |
| Sonnet 5.5, xhigh | 52.1% | 64.4% |
At max effort, Haiku 5.5 edges Sonnet 5.5 on Main, 46.4% to 46.2%. At each model’s best setting, Sonnet leads by 5.7 points. Say which effort you mean when you quote either.
System card extras
The system card adds rows the launch post skipped. All are max effort unless noted.
| Benchmark | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| SWE-Bench Pro | 64.8 | n/r | 81.3 |
| SWE-bench Multilingual | 83.7 | 67.4 | 90.3 |
| SWE-bench Multimodal | 30.7 | 19.8 | 54.3 |
| OSWorld 2.1, strict pass rate | 37.1% | n/r | 48.8% |
| Chartography, with tools | 86.2% | 8.8% | 90.2% |
| OfficeQA / OfficeQA Pro | 73.5% / 60.3% | 63.0% / 47.1% | 76.9% / 65.6% |
| HealthBench Professional (length-adjusted) | 64.8% | 32.2% | 69.2% |
OSWorld’s 72.4% is partial credit; only 37.1% of tasks pass outright. Chartography nearly doubles with tools (46.4% to 86.2%), so give Haiku 5.5 tools for chart work.
Default effort scores lower
Haiku 5.5 defaults to medium on the Claude API and in Claude Code. The system card reports these default-effort results:
| Benchmark | Medium (default) | Max | Note |
|---|---|---|---|
| GDPval-AA v2.1 | 1277 | 1620 | Medium used about a tenth of max’s output tokens |
| AA-Briefcase v1.1 | 1372 | 1578 | Medium used under a quarter of max’s output tokens |
| HealthBench Professional | 59.9% | 64.8% | Low 57.9%, high 61.3% |
Call the API without output_config.effort and expect the left column. The effort docs cover all five levels.
Score and cost by effort
From the launch charts, as score (cost). OSWorld and Terminal-Bench cost is per attempt; GDPval-AA is per task.
| Model | Low | Medium | High | Xhigh | Max |
|---|---|---|---|---|---|
| OSWorld 2.1 | |||||
| Haiku 5.5 | 42.0% ($0.07) | 53.3% ($0.13) | 61.3% ($0.18) | 67.6% ($0.28) | 72.4% ($0.61) |
| Sonnet 5.5 | 57.9% ($0.68) | 66.0% ($0.93) | 73.2% ($1.38) | 81.1% ($2.22) | 83.9% ($5.73) |
| GPT-6 Luna | 19.2% ($0.04) | 37.5% ($0.13) | 42.3% ($0.14) | 44.8% ($0.17) | 48.9% ($0.21) |
| GDPval-AA v2.1 | |||||
| Haiku 5.5 | 1125 ($0.012) | 1277 ($0.030) | 1420 ($0.089) | 1513 ($0.27) | 1620 ($0.87) |
| Sonnet 5.5 | 1179 ($0.22) | 1324 ($0.27) | 1551 ($0.62) | 1731 ($1.88) | 1840 ($6.78) |
| GPT-6 Luna | 1036 ($0.004) | 1262 ($0.02) | 1344 ($0.03) | 1364 ($0.05) | 1437 ($0.09) |
| Terminal-Bench 4.0 | |||||
| Haiku 5.5 | 12.7% ($0.42) | 20.3% ($0.68) | 24.8% ($1.04) | 31.5% ($1.75) | 39.2% ($2.64) |
| Sonnet 5.5 | 20.0% ($0.62) | 28.8% ($0.68) | 43.0% ($1.46) | 61.5% ($4.34) | 70.6% ($10.44) |
- Max is expensive for its last step. On GDPval-AA, max costs about 28x medium for 343 more Elo points. On OSWorld, max costs over twice xhigh for 4.8 points.
- Haiku 5.5 at medium beats Luna at max on OSWorld: 53.3% for $0.13 versus 48.9% for $0.21. Luna is cheaper at every other matching level, though; at medium the two cost about the same. See Haiku 5.5 vs GPT-6 Luna.
- Haiku 5.5 at max beats Sonnet 5.5 at medium on OSWorld (72.4% for $0.61 versus 66.0% for $0.93).
- Terminal-Bench is the exception. Sonnet 5.5 at high (43.0%, $1.46) beats Haiku 5.5 at max (39.2%, $2.64) for less money. Sonnet costs use the new $0.10 cache-read price.
Costs use list prices: $0.10/$0.50 per million tokens for prompts up to 100K tokens, higher above that (see the pricing breakdown).
Where Haiku 5.5 still trails
Sonnet 5.5 leads every row of the launch table. Haiku’s only head-to-head win is FrontierCode at matched max effort. The widest gap is Terminal-Bench 4.0: 39.2% versus 70.6%. SWE-Bench Pro (64.8 versus 81.3) and SWE-bench Multimodal (30.7 versus 54.3) show the same pattern.
Anthropic agrees: Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks.” Haiku 5.5 fits narrowly scoped work: compaction, summarization, classification, browser use and subagents under a larger model. Running it as a cheap subagent is covered in Haiku 5.5 in Claude Code.
Independent results: none yet
As of October 8, 2026, no independent lab has published Haiku 5.5 results. Artificial Analysis’s Haiku 5.5 page returns a 404, and its leaderboard lists only Claude 4.5 Haiku. Vals, LMArena, SWE-bench and Aider showed nothing either. There’s no third-party speed figure; Anthropic calls Haiku 5.5 its “fastest model to date” at standard speed, slower than Opus in Fast Mode.
The closest outside number comes from Cursor. Its model docs say Haiku 5.5 “scores 48.4% on CursorBench at max effort,” up from 30.9% at low. With thinking off, Cursor reports 22.1% to 26.2% at the same effort levels where thinking on scored 30.9% to 42.3%.
What customers report
These come from Anthropic’s launch post and were not independently verified:
- HubSpot: 92.8% averaged over three runs on its simulated CRM portal suite, the best score it has seen on that suite.
- AlphaSense: 0.84 versus Haiku 4.5’s 0.76 on 400 Ask in Document queries, which it calls statistically significant.
- Box: 11 points higher than Haiku 4.5 at about half the latency, in early testing.
- Asana: over 30% lower latency for task completions versus its current model.
- Cognition: Devin Fusion holds a FrontierCode score of 66.2 with Haiku 5.5 as sidekick and Opus 5.5 as lead. That’s a two-model system, not Haiku alone.
Run your own eval
Benchmarks don’t look like your traffic. Anthropic’s prompting guide suggests using xhigh or max only where your evals show a gain, and running the same evals on Sonnet 5.5 to compare. In Apidog:
- Store
ANTHROPIC_API_KEYas an environment variable and save 20 to 50 real prompts as Messages requests. - Add assertions: status 200, a
stop_reasonother than"max_tokens"or"refusal", and the content a correct answer needs. - Run the set at
low,mediumandhigh. Changing effort invalidates the prompt cache. - Record pass rate and the token counts in
usage. The new tokenizer produces about 30% more tokens than Haiku 4.5’s for the same text. - Duplicate the collection, swap in
claude-sonnet-5-5, and compare both models in one project.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"max_tokens": 8000,
"thinking": {"type": "adaptive"},
"output_config": {"effort": "medium"},
"messages": [{"role": "user", "content": "Classify this support ticket as billing, bug or feature request: ..."}]
}'
Don’t send temperature, top_p, top_k or an assistant prefill; each returns a 400 on Haiku 5.5. Select response blocks by type, since a thinking block can come first. See the API guide and testing LLM applications.
FAQ
Is Claude Haiku 5.5 better than Sonnet 5.5? Not on Anthropic’s numbers. Sonnet 5.5 leads every launch-table row; Haiku only edges it on FrontierCode when both run at max (46.4% vs 46.2%). See Haiku 5.5 vs Haiku 4.5 for the upgrade view.
What is Haiku 5.5’s SWE-bench score? 64.8 on SWE-Bench Pro, 83.7 on SWE-bench Multilingual and 30.7 on SWE-bench Multimodal, all at max effort.
What effort were the benchmarks run at? Max, usually averaged over five trials. The API default is medium, where GDPval-AA scores 1277 instead of 1620.
Are there independent Haiku 5.5 benchmarks? Not yet. Artificial Analysis ran GDPval-AA and AA-Briefcase independently, but Anthropic published those scores; AA has no Haiku 5.5 page.
Is GPT-6 Luna cheaper? Usually. Luna costs less at every GDPval-AA effort level and every OSWorld level except medium, where the two cost about the same. Haiku 5.5 scores higher on both.
Next step
Start at medium and step up only where your pass rate gain pays for itself. Download Apidog, build the eval collection above, and keep the results in Apidog next to your Sonnet 5.5 baseline before you switch production traffic.



