Claude Haiku 5.5 vs GPT-6 Luna

Haiku 5.5 vs GPT-6 Luna: same $0.10/$0.50 price under 100K tokens, different long-prompt tiers, Anthropic's benchmarks, and per-effort cost data.

Medy Evrard

8 October 2026

Claude Haiku 5.5 vs GPT-6 Luna

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Claude Haiku 5.5 and GPT-6 Luna share a list price: $0.10/$0.50 per million input/output tokens for Haiku prompts up to 100K tokens, with $0.01 cache reads and $0.125 cache writes on both. The difference is where that price ends. Haiku 5.5 moves to $0.50/$2.50 over 100K tokens; Luna holds its base rate until 272K input tokens, then charges 2x input and 1.5x output. On Anthropic’s launch table, Haiku 5.5 scores higher on every shared row. On Anthropic’s own cost charts, Luna is cheaper per task at every GDPval-AA effort level and at four of five on OSWorld; at OSWorld medium, Haiku 5.5 is marginally cheaper.

button

Below: price, specs, the shared benchmarks and who ran them, per-effort cost curves, and routing by workload. For each model alone, see what is Claude Haiku 5.5 and what is GPT-6 Luna. The last section sends the same prompt to both in Apidog.

Price side by side

Per 1M tokens Claude Haiku 5.5 GPT-6 Luna
Input, base tier $0.10 (prompts up to 100K tokens) $0.10 (up to 272K input tokens)
Output, base tier $0.50 $0.50
Cache read, base tier $0.01 $0.01
Cache write, base tier $0.125 (5-minute) $0.125
Higher tier, input / output $0.50 / $2.50 (over 100K) $0.20 / $0.75 (over 272K)
Higher tier, cache read / write $0.05 / $0.625 2x cache rates
Batch $0.05 / $0.25 up to 100K; $0.25 / $1.25 over 50% of Standard rates (Batch and Flex)

Sources: Anthropic’s pricing docs and OpenAI’s GPT-6 Luna model page. OpenAI reprices “the full request” above 272K. Anthropic says “a prompt of over 100,000 tokens pays higher prices,” charged per request by prompt length.

Below 100K tokens, the bills match line for line. Between 100K and 272K, Haiku 5.5 is 5x Luna on input and output. Above 272K, both are in their higher tier, and Luna’s is still cheaper.

Two examples, counted in each vendor’s own tokens:

Anthropic says prompts up to 100K tokens made up around 90% of Haiku 4.5 requests. If your traffic looks like that, most of your calls cost the same on either model. More math in Claude Haiku 5.5 pricing.

Tokenizers move the threshold

A token on one model isn’t a token on the other. Haiku 5.5’s updated tokenizer produces about 30% more tokens than Haiku 4.5 for the same text, per Anthropic. In the Hacker News launch thread, one commenter estimated that “modern Claude’s 100K tokens are about ~60-65K modern GPT tokens.” That’s a community estimate, not a vendor figure. If it holds, Haiku’s 100K line arrives sooner in text terms than the raw numbers suggest. Count your own prompts on both APIs and read each usage object.

Specs

Claude Haiku 5.5 GPT-6 Luna
API model ID claude-haiku-5-5 gpt-6-luna
Context window 1M tokens 1,050,000 tokens
Max output 128K tokens (300K on Batch with a beta header) 128,000 tokens
Knowledge cutoff June 2026 May 18, 2026
Effort levels low, medium (default), high, xhigh, max none, low, medium (default), high, xhigh, max
Thinking off thinking: {"type": "disabled"} at high effort or below none reasoning effort
Free chat access Free Claude.ai users can select it Free and Go users in the ChatGPT desktop app

Both default to medium. On Haiku 5.5 the parameter is output_config.effort; on Luna it’s reasoning.effort. Neither free chat route is an API key, so it can’t drive a script, CI job or agent. See Claude Haiku 5.5 for free and GPT-6 Luna for free.

One caching difference matters in loops. On Haiku 5.5, changing top-level effort between requests invalidates the prompt cache. OpenAI says changing reasoning effort no longer breaks the cache on GPT-6 models (GPT-6 prompt caching).

Benchmarks: Anthropic’s table, and who ran each Luna cell

Every number below comes from Anthropic’s launch post and its system card. Anthropic compiled the table, but different parties ran different rows. Haiku 5.5 results are at max effort, mostly averaged over five trials.

Benchmark Haiku 5.5 GPT-6 Luna How the Luna cell was sourced
GDPval-AA v2.1 (Elo) 1620 1437 Run independently by Artificial Analysis
AA-Briefcase v1.1 (Elo) 1578 1336 Run independently by Artificial Analysis
OSWorld 2.1, offline subset 72.4% 48.9% Run by Anthropic on the same 82 tasks via OpenAI’s API
Terminal-Bench 4.0 39.2% 16.4% Public leaderboard, Codex CLI at max effort
FrontierCode 1.1 (Main) 46.4% 42.4% Run by Cognition (Claude in Claude Code, GPT in Codex)
Chartography, no tools 46.4% 29.1% As publicly reported by Surge AI

The rows share a benchmark name, not a harness: Terminal-Bench pits Haiku 5.5 in Claude Code against Luna in the Codex CLI. OSWorld is the closest to controlled, since Anthropic ran both models on the same tasks, but it’s still one vendor running a rival’s model.

No independent head-to-head exists yet. As of October 8, 2026, Artificial Analysis has no Haiku 5.5 page, and we found no Haiku 5.5 result on Vals, LMArena, SWE-bench or Aider. AA’s leaderboard that day lists GPT-6 Luna (max) at an Intelligence Index of 38 and 129 output tokens per second. Full Haiku 5.5 results: Claude Haiku 5.5 benchmarks.

Cost per attempt at each effort level

The launch table shows each model at max. Anthropic’s launch page also plots each effort level against cost, with Luna on two charts.

OSWorld 2.1 (offline subset), partial-credit score / cost per attempt:

Effort Haiku 5.5 GPT-6 Luna
Low 42.0% / $0.0695 19.2% / $0.038
Medium 53.3% / $0.1257 37.5% / $0.1261
High 61.3% / $0.1827 42.3% / $0.1435
Xhigh 67.6% / $0.2792 44.8% / $0.1672
Max 72.4% / $0.6111 48.9% / $0.205

GDPval-AA v2.1, Elo / cost per task:

Effort Haiku 5.5 GPT-6 Luna
Low 1125 / $0.01167 1036 / $0.0038
Medium 1277 / $0.03042 1262 / $0.02
High 1420 / $0.08938 1344 / $0.03
Xhigh 1513 / $0.26725 1364 / $0.05
Max 1620 / $0.86593 1437 / $0.09

Luna costs less at every effort level on GDPval-AA and at low, high, xhigh and max on OSWorld. At OSWorld medium, Haiku 5.5 is marginally cheaper ($0.1257 vs $0.1261). Haiku 5.5 scores higher at every level. Compare at matched spend:

So on computer use, Haiku 5.5 buys more score per dollar in Anthropic’s run. On knowledge work, the two are close at equal spend, and Luna’s lower floor wins when good-enough is the goal.

Routing: which model for which workload

Anthropic positions Haiku 5.5 for “narrowly scoped tasks” and says Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks.” Both models here are built for volume.

Workload Start with Why
Classification, routing, extraction under 100K tokens Either; test both Identical price; pick on accuracy for your labels
Long-document Q&A, 100K to 272K tokens GPT-6 Luna Base rate holds to 272K; Haiku 5.5 is 5x above 100K
Browser and desktop automation Claude Haiku 5.5 Higher OSWorld score at matched cost
Subagents under Opus 5.5 or Sonnet 5.5 Claude Haiku 5.5 Claude Code model: haiku frontmatter (Haiku 5.5 on the Anthropic API)
Lowest cost per call GPT-6 Luna Cheapest floor: $0.038 (OSWorld) and $0.0038 (GDPval-AA) at low
Knowledge-work drafting at high effort Either; test both Within 17 Elo at matched spend on GDPval-AA

See Claude Haiku 5.5 in Claude Code for subagents and GPT-6 Luna for high-volume API workloads for high-QPS pipelines. Coming from Haiku 4.5? Haiku 5.5 vs Haiku 4.5 lists the five new 400 errors.

Test both side by side in Apidog

The head-to-head that settles it runs on your prompts. In Apidog, create one project with two requests and the same prompt:

  1. Store ANTHROPIC_API_KEY and your OpenAI key as environment variables.
  2. Add the Haiku 5.5 request below, then the Luna request with the same prompt and effort.
  3. Assert that Haiku’s stop_reason isn’t "refusal" (there’s no server-side fallback) and that usage stays under your token budget.
  4. Run both at low, medium and high effort and save each response to compare tokens, latency and output.
curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-haiku-5-5",
    "max_tokens": 2048,
    "thinking": {"type": "adaptive"},
    "output_config": {"effort": "medium"},
    "messages": [
      {"role": "user", "content": "Classify this support ticket as billing, bug, or feature request: The export button returns a 500 error since this morning."}
    ]
  }'

Leave out temperature, top_p and top_k: non-default values return a 400 on Haiku 5.5. Walkthrough: how to use the Claude Haiku 5.5 API. For eval sets, see testing LLM applications.

FAQ

Are Claude Haiku 5.5 and GPT-6 Luna the same price? Yes, for Haiku prompts up to 100K tokens: $0.10/$0.50, $0.01 cache reads, $0.125 cache writes. Above that, Haiku 5.5 costs $0.50/$2.50 while Luna stays at base until 272K.

Which model scores higher? Haiku 5.5, on every shared row of Anthropic’s launch table. Anthropic compiled that table; no independent head-to-head exists yet.

Which is cheaper per task? Usually Luna: at every effort level on Anthropic’s GDPval-AA chart and four of five on OSWorld (at medium, Haiku 5.5 is $0.0004 cheaper). At matched spend, Haiku 5.5 leads on OSWorld; GDPval-AA is close.

Is either one free? In chat only: Free Claude.ai users can select Haiku 5.5, and Free and Go users get Luna in the ChatGPT desktop app. Neither gives you an API key. See Claude Haiku 5.5 for free.

Pick by prompt length, then test

Check your prompt lengths first. Under 100K tokens, price is a tie and accuracy decides, where Anthropic’s numbers favor Haiku 5.5. Between 100K and 272K, Luna’s threshold saves money on every call. Then run both on real requests at two or three effort levels: download Apidog, put both requests in one project, and let the usage numbers decide.

Explore more

Claude Haiku 5.5 Benchmarks

Claude Haiku 5.5 Benchmarks

Claude Haiku 5.5 benchmarks: 72.4% OSWorld, 1620 GDPval-AA, 39.2% Terminal-Bench at max effort. Who ran each test, per-effort costs, and your own eval.

8 October 2026

Claude Haiku 5.5 vs Haiku 4.5: What Changed and the Breaking Changes to Fix First

Claude Haiku 5.5 vs Haiku 4.5: What Changed and the Breaking Changes to Fix First

Haiku 5.5 vs Haiku 4.5: 90% cheaper up to 100K tokens, 1M context, and five breaking changes that return 400s. Before/after JSON fixes inside.

8 October 2026

Claude Haiku 5.5 Pricing

Claude Haiku 5.5 Pricing

Claude Haiku 5.5 pricing: $0.10/$0.50 per million tokens for prompts up to 100K, $0.50/$2.50 above. Caching, batch, and worked cost examples.

8 October 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Claude Haiku 5.5 vs GPT-6 Luna