Claude Haiku 5.5 Benchmarks

Claude Haiku 5.5 benchmarks: 72.4% OSWorld, 1620 GDPval-AA, 39.2% Terminal-Bench at max effort. Who ran each test, per-effort costs, and your own eval.

Ashley Goolam

Ashley Goolam

8 October 2026

Claude Haiku 5.5 Benchmarks

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Claude Haiku 5.5 scores 72.4% on OSWorld 2.1, 1620 on GDPval-AA v2.1, 46.4% on FrontierCode 1.1 and 39.2% on Terminal-Bench 4.0. Haiku 4.5 scored 15.7%, 735 and 0.0% on three of those. All are max-effort numbers; at the default medium effort, GDPval-AA drops to 1277. No independent lab has published Haiku 5.5 results yet.

Below: every Claude Haiku 5.5 benchmark, who ran it, how effort moves score and cost, and how to test the model on your own traffic in Apidog. For specs and pricing, start with what is Claude Haiku 5.5.

button

The launch table

Every Haiku 5.5 cell uses adaptive thinking at max effort, usually averaged over five trials (Terminal-Bench used ten per task). “n/r” means no result was published.

Benchmark Haiku 5.5 Haiku 4.5 GPT-6 Luna Sonnet 5.5
GDPval-AA v2.1 (Elo) 1620 735 1437 1840
AA-Briefcase v1.1 (Elo) 1578 614 1336 1824
OSWorld 2.1, offline subset 72.4% 15.7% 48.9% 83.9%
Humanity’s Last Exam, no tools 45.9% 10.2% n/r 56.9%
Humanity’s Last Exam, with tools 57.4% 18.7% n/r 64.5%
Terminal-Bench 4.0 39.2% 0.0% 16.4% 70.6%
FrontierCode 1.1 (Main) 46.4% n/r 42.4% 52.1% (xhigh)
Chartography, no tools 46.4% 6.4% 29.1% 61.6%

Who ran each benchmark

Anthropic ran most tests itself. Cross-vendor rows share a benchmark name, not always a harness or effort setting.

Benchmark Who ran it Conditions
GDPval-AA v2.1, AA-Briefcase v1.1 Artificial Analysis, independently GDPval-AA: 220 tasks, 44 occupations, Elo anchored to DeepSeek V4.1 Flash (max) at 1600; Briefcase: multi-week projects
FrontierCode 1.1 Cognition Claude in Claude Code, GPT in Codex; Main = hardest 100 of 150 tasks
Chartography Anthropic, on Surge AI’s benchmark Gemini 3.5 Flash grader; Luna score from Surge AI
Terminal-Bench 4.0 Anthropic Claude Code --bare, 66 tasks, 10 trials each, safeguards on
OSWorld 2.1 Anthropic 82 of 108 tasks, 1080p, up to 500 steps
Humanity’s Last Exam Anthropic 980K-token task budget, Opus 4.6 grader

Two GPT-6 Luna cells need context. Anthropic ran Luna’s OSWorld score through OpenAI’s API on the same 82 tasks. Luna’s Terminal-Bench 16.4% comes from the public leaderboard (Codex CLI, max effort). On Terminal-Bench, safeguards stopped 1.8% of Haiku 5.5 trials (12 of 660), and each counted as a fail.

The FrontierCode trap

The launch table puts Haiku 5.5 at max (46.4%) next to Sonnet 5.5 at xhigh (52.1%), Sonnet’s best score. The system card gives the like-for-like numbers:

FrontierCode 1.1 Main Extended
Haiku 5.5, max 46.4% 58.4%
Haiku 5.5, xhigh 45.8% n/r
Sonnet 5.5, max 46.2% 59.1%
Sonnet 5.5, xhigh 52.1% 64.4%

At max effort, Haiku 5.5 edges Sonnet 5.5 on Main, 46.4% to 46.2%. At each model’s best setting, Sonnet leads by 5.7 points. Say which effort you mean when you quote either.

System card extras

The system card adds rows the launch post skipped. All are max effort unless noted.

Benchmark Haiku 5.5 Haiku 4.5 Sonnet 5.5
SWE-Bench Pro 64.8 n/r 81.3
SWE-bench Multilingual 83.7 67.4 90.3
SWE-bench Multimodal 30.7 19.8 54.3
OSWorld 2.1, strict pass rate 37.1% n/r 48.8%
Chartography, with tools 86.2% 8.8% 90.2%
OfficeQA / OfficeQA Pro 73.5% / 60.3% 63.0% / 47.1% 76.9% / 65.6%
HealthBench Professional (length-adjusted) 64.8% 32.2% 69.2%

OSWorld’s 72.4% is partial credit; only 37.1% of tasks pass outright. Chartography nearly doubles with tools (46.4% to 86.2%), so give Haiku 5.5 tools for chart work.

Default effort scores lower

Haiku 5.5 defaults to medium on the Claude API and in Claude Code. The system card reports these default-effort results:

Benchmark Medium (default) Max Note
GDPval-AA v2.1 1277 1620 Medium used about a tenth of max’s output tokens
AA-Briefcase v1.1 1372 1578 Medium used under a quarter of max’s output tokens
HealthBench Professional 59.9% 64.8% Low 57.9%, high 61.3%

Call the API without output_config.effort and expect the left column. The effort docs cover all five levels.

Score and cost by effort

From the launch charts, as score (cost). OSWorld and Terminal-Bench cost is per attempt; GDPval-AA is per task.

Model Low Medium High Xhigh Max
OSWorld 2.1
Haiku 5.5 42.0% ($0.07) 53.3% ($0.13) 61.3% ($0.18) 67.6% ($0.28) 72.4% ($0.61)
Sonnet 5.5 57.9% ($0.68) 66.0% ($0.93) 73.2% ($1.38) 81.1% ($2.22) 83.9% ($5.73)
GPT-6 Luna 19.2% ($0.04) 37.5% ($0.13) 42.3% ($0.14) 44.8% ($0.17) 48.9% ($0.21)
GDPval-AA v2.1
Haiku 5.5 1125 ($0.012) 1277 ($0.030) 1420 ($0.089) 1513 ($0.27) 1620 ($0.87)
Sonnet 5.5 1179 ($0.22) 1324 ($0.27) 1551 ($0.62) 1731 ($1.88) 1840 ($6.78)
GPT-6 Luna 1036 ($0.004) 1262 ($0.02) 1344 ($0.03) 1364 ($0.05) 1437 ($0.09)
Terminal-Bench 4.0
Haiku 5.5 12.7% ($0.42) 20.3% ($0.68) 24.8% ($1.04) 31.5% ($1.75) 39.2% ($2.64)
Sonnet 5.5 20.0% ($0.62) 28.8% ($0.68) 43.0% ($1.46) 61.5% ($4.34) 70.6% ($10.44)

Costs use list prices: $0.10/$0.50 per million tokens for prompts up to 100K tokens, higher above that (see the pricing breakdown).

Where Haiku 5.5 still trails

Sonnet 5.5 leads every row of the launch table. Haiku’s only head-to-head win is FrontierCode at matched max effort. The widest gap is Terminal-Bench 4.0: 39.2% versus 70.6%. SWE-Bench Pro (64.8 versus 81.3) and SWE-bench Multimodal (30.7 versus 54.3) show the same pattern.

Anthropic agrees: Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks.” Haiku 5.5 fits narrowly scoped work: compaction, summarization, classification, browser use and subagents under a larger model. Running it as a cheap subagent is covered in Haiku 5.5 in Claude Code.

Independent results: none yet

As of October 8, 2026, no independent lab has published Haiku 5.5 results. Artificial Analysis’s Haiku 5.5 page returns a 404, and its leaderboard lists only Claude 4.5 Haiku. Vals, LMArena, SWE-bench and Aider showed nothing either. There’s no third-party speed figure; Anthropic calls Haiku 5.5 its “fastest model to date” at standard speed, slower than Opus in Fast Mode.

The closest outside number comes from Cursor. Its model docs say Haiku 5.5 “scores 48.4% on CursorBench at max effort,” up from 30.9% at low. With thinking off, Cursor reports 22.1% to 26.2% at the same effort levels where thinking on scored 30.9% to 42.3%.

What customers report

These come from Anthropic’s launch post and were not independently verified:

Run your own eval

Benchmarks don’t look like your traffic. Anthropic’s prompting guide suggests using xhigh or max only where your evals show a gain, and running the same evals on Sonnet 5.5 to compare. In Apidog:

  1. Store ANTHROPIC_API_KEY as an environment variable and save 20 to 50 real prompts as Messages requests.
  2. Add assertions: status 200, a stop_reason other than "max_tokens" or "refusal", and the content a correct answer needs.
  3. Run the set at low, medium and high. Changing effort invalidates the prompt cache.
  4. Record pass rate and the token counts in usage. The new tokenizer produces about 30% more tokens than Haiku 4.5’s for the same text.
  5. Duplicate the collection, swap in claude-sonnet-5-5, and compare both models in one project.
curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-haiku-5-5",
    "max_tokens": 8000,
    "thinking": {"type": "adaptive"},
    "output_config": {"effort": "medium"},
    "messages": [{"role": "user", "content": "Classify this support ticket as billing, bug or feature request: ..."}]
  }'

Don’t send temperature, top_p, top_k or an assistant prefill; each returns a 400 on Haiku 5.5. Select response blocks by type, since a thinking block can come first. See the API guide and testing LLM applications.

FAQ

Is Claude Haiku 5.5 better than Sonnet 5.5? Not on Anthropic’s numbers. Sonnet 5.5 leads every launch-table row; Haiku only edges it on FrontierCode when both run at max (46.4% vs 46.2%). See Haiku 5.5 vs Haiku 4.5 for the upgrade view.

What is Haiku 5.5’s SWE-bench score? 64.8 on SWE-Bench Pro, 83.7 on SWE-bench Multilingual and 30.7 on SWE-bench Multimodal, all at max effort.

What effort were the benchmarks run at? Max, usually averaged over five trials. The API default is medium, where GDPval-AA scores 1277 instead of 1620.

Are there independent Haiku 5.5 benchmarks? Not yet. Artificial Analysis ran GDPval-AA and AA-Briefcase independently, but Anthropic published those scores; AA has no Haiku 5.5 page.

Is GPT-6 Luna cheaper? Usually. Luna costs less at every GDPval-AA effort level and every OSWorld level except medium, where the two cost about the same. Haiku 5.5 scores higher on both.

Next step

Start at medium and step up only where your pass rate gain pays for itself. Download Apidog, build the eval collection above, and keep the results in Apidog next to your Sonnet 5.5 baseline before you switch production traffic.

Explore more

Claude Haiku 5.5 vs GPT-6 Luna

Claude Haiku 5.5 vs GPT-6 Luna

Haiku 5.5 vs GPT-6 Luna: same $0.10/$0.50 price under 100K tokens, different long-prompt tiers, Anthropic's benchmarks, and per-effort cost data.

8 October 2026

Claude Haiku 5.5 vs Haiku 4.5: What Changed and the Breaking Changes to Fix First

Claude Haiku 5.5 vs Haiku 4.5: What Changed and the Breaking Changes to Fix First

Haiku 5.5 vs Haiku 4.5: 90% cheaper up to 100K tokens, 1M context, and five breaking changes that return 400s. Before/after JSON fixes inside.

8 October 2026

Claude Haiku 5.5 Pricing

Claude Haiku 5.5 Pricing

Claude Haiku 5.5 pricing: $0.10/$0.50 per million tokens for prompts up to 100K, $0.50/$2.50 above. Caching, batch, and worked cost examples.

8 October 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Claude Haiku 5.5 Benchmarks