Claude Haiku 5.5 and GPT-6 Luna share a list price: $0.10/$0.50 per million input/output tokens for Haiku prompts up to 100K tokens, with $0.01 cache reads and $0.125 cache writes on both. The difference is where that price ends. Haiku 5.5 moves to $0.50/$2.50 over 100K tokens; Luna holds its base rate until 272K input tokens, then charges 2x input and 1.5x output. On Anthropic’s launch table, Haiku 5.5 scores higher on every shared row. On Anthropic’s own cost charts, Luna is cheaper per task at every GDPval-AA effort level and at four of five on OSWorld; at OSWorld medium, Haiku 5.5 is marginally cheaper.
Below: price, specs, the shared benchmarks and who ran them, per-effort cost curves, and routing by workload. For each model alone, see what is Claude Haiku 5.5 and what is GPT-6 Luna. The last section sends the same prompt to both in Apidog.
Price side by side
| Per 1M tokens | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Input, base tier | $0.10 (prompts up to 100K tokens) | $0.10 (up to 272K input tokens) |
| Output, base tier | $0.50 | $0.50 |
| Cache read, base tier | $0.01 | $0.01 |
| Cache write, base tier | $0.125 (5-minute) | $0.125 |
| Higher tier, input / output | $0.50 / $2.50 (over 100K) | $0.20 / $0.75 (over 272K) |
| Higher tier, cache read / write | $0.05 / $0.625 | 2x cache rates |
| Batch | $0.05 / $0.25 up to 100K; $0.25 / $1.25 over | 50% of Standard rates (Batch and Flex) |
Sources: Anthropic’s pricing docs and OpenAI’s GPT-6 Luna model page. OpenAI reprices “the full request” above 272K. Anthropic says “a prompt of over 100,000 tokens pays higher prices,” charged per request by prompt length.
Below 100K tokens, the bills match line for line. Between 100K and 272K, Haiku 5.5 is 5x Luna on input and output. Above 272K, both are in their higher tier, and Luna’s is still cheaper.
Two examples, counted in each vendor’s own tokens:
- 150K-token prompt, 2K output. Haiku 5.5: 0.15 × $0.50 + 0.002 × $2.50 = $0.080. Luna: 0.15 × $0.10 + 0.002 × $0.50 = $0.016.
- 300K-token prompt, 2K output. Haiku 5.5: 0.30 × $0.50 + 0.002 × $2.50 = $0.155. Luna: 0.30 × $0.20 + 0.002 × $0.75 = $0.0615.
Anthropic says prompts up to 100K tokens made up around 90% of Haiku 4.5 requests. If your traffic looks like that, most of your calls cost the same on either model. More math in Claude Haiku 5.5 pricing.
Tokenizers move the threshold
A token on one model isn’t a token on the other. Haiku 5.5’s updated tokenizer produces about 30% more tokens than Haiku 4.5 for the same text, per Anthropic. In the Hacker News launch thread, one commenter estimated that “modern Claude’s 100K tokens are about ~60-65K modern GPT tokens.” That’s a community estimate, not a vendor figure. If it holds, Haiku’s 100K line arrives sooner in text terms than the raw numbers suggest. Count your own prompts on both APIs and read each usage object.
Specs
| Claude Haiku 5.5 | GPT-6 Luna | |
|---|---|---|
| API model ID | claude-haiku-5-5 |
gpt-6-luna |
| Context window | 1M tokens | 1,050,000 tokens |
| Max output | 128K tokens (300K on Batch with a beta header) | 128,000 tokens |
| Knowledge cutoff | June 2026 | May 18, 2026 |
| Effort levels | low, medium (default), high, xhigh, max | none, low, medium (default), high, xhigh, max |
| Thinking off | thinking: {"type": "disabled"} at high effort or below |
none reasoning effort |
| Free chat access | Free Claude.ai users can select it | Free and Go users in the ChatGPT desktop app |
Both default to medium. On Haiku 5.5 the parameter is output_config.effort; on Luna it’s reasoning.effort. Neither free chat route is an API key, so it can’t drive a script, CI job or agent. See Claude Haiku 5.5 for free and GPT-6 Luna for free.
One caching difference matters in loops. On Haiku 5.5, changing top-level effort between requests invalidates the prompt cache. OpenAI says changing reasoning effort no longer breaks the cache on GPT-6 models (GPT-6 prompt caching).
Benchmarks: Anthropic’s table, and who ran each Luna cell
Every number below comes from Anthropic’s launch post and its system card. Anthropic compiled the table, but different parties ran different rows. Haiku 5.5 results are at max effort, mostly averaged over five trials.
| Benchmark | Haiku 5.5 | GPT-6 Luna | How the Luna cell was sourced |
|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 1437 | Run independently by Artificial Analysis |
| AA-Briefcase v1.1 (Elo) | 1578 | 1336 | Run independently by Artificial Analysis |
| OSWorld 2.1, offline subset | 72.4% | 48.9% | Run by Anthropic on the same 82 tasks via OpenAI’s API |
| Terminal-Bench 4.0 | 39.2% | 16.4% | Public leaderboard, Codex CLI at max effort |
| FrontierCode 1.1 (Main) | 46.4% | 42.4% | Run by Cognition (Claude in Claude Code, GPT in Codex) |
| Chartography, no tools | 46.4% | 29.1% | As publicly reported by Surge AI |
The rows share a benchmark name, not a harness: Terminal-Bench pits Haiku 5.5 in Claude Code against Luna in the Codex CLI. OSWorld is the closest to controlled, since Anthropic ran both models on the same tasks, but it’s still one vendor running a rival’s model.
No independent head-to-head exists yet. As of October 8, 2026, Artificial Analysis has no Haiku 5.5 page, and we found no Haiku 5.5 result on Vals, LMArena, SWE-bench or Aider. AA’s leaderboard that day lists GPT-6 Luna (max) at an Intelligence Index of 38 and 129 output tokens per second. Full Haiku 5.5 results: Claude Haiku 5.5 benchmarks.
Cost per attempt at each effort level
The launch table shows each model at max. Anthropic’s launch page also plots each effort level against cost, with Luna on two charts.
OSWorld 2.1 (offline subset), partial-credit score / cost per attempt:
| Effort | Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Low | 42.0% / $0.0695 | 19.2% / $0.038 |
| Medium | 53.3% / $0.1257 | 37.5% / $0.1261 |
| High | 61.3% / $0.1827 | 42.3% / $0.1435 |
| Xhigh | 67.6% / $0.2792 | 44.8% / $0.1672 |
| Max | 72.4% / $0.6111 | 48.9% / $0.205 |
GDPval-AA v2.1, Elo / cost per task:
| Effort | Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Low | 1125 / $0.01167 | 1036 / $0.0038 |
| Medium | 1277 / $0.03042 | 1262 / $0.02 |
| High | 1420 / $0.08938 | 1344 / $0.03 |
| Xhigh | 1513 / $0.26725 | 1364 / $0.05 |
| Max | 1620 / $0.86593 | 1437 / $0.09 |
Luna costs less at every effort level on GDPval-AA and at low, high, xhigh and max on OSWorld. At OSWorld medium, Haiku 5.5 is marginally cheaper ($0.1257 vs $0.1261). Haiku 5.5 scores higher at every level. Compare at matched spend:
- OSWorld: at medium, Haiku 5.5 costs $0.1257 per attempt to Luna’s $0.1261 and scores 53.3% to Luna’s 37.5%. Haiku at medium also beats Luna at max (48.9%) for less money.
- GDPval-AA: Luna at max (1437, $0.09) and Haiku 5.5 at high (1420, $0.08938) cost about the same and land 17 Elo apart.
So on computer use, Haiku 5.5 buys more score per dollar in Anthropic’s run. On knowledge work, the two are close at equal spend, and Luna’s lower floor wins when good-enough is the goal.
Routing: which model for which workload
Anthropic positions Haiku 5.5 for “narrowly scoped tasks” and says Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks.” Both models here are built for volume.
| Workload | Start with | Why |
|---|---|---|
| Classification, routing, extraction under 100K tokens | Either; test both | Identical price; pick on accuracy for your labels |
| Long-document Q&A, 100K to 272K tokens | GPT-6 Luna | Base rate holds to 272K; Haiku 5.5 is 5x above 100K |
| Browser and desktop automation | Claude Haiku 5.5 | Higher OSWorld score at matched cost |
| Subagents under Opus 5.5 or Sonnet 5.5 | Claude Haiku 5.5 | Claude Code model: haiku frontmatter (Haiku 5.5 on the Anthropic API) |
| Lowest cost per call | GPT-6 Luna | Cheapest floor: $0.038 (OSWorld) and $0.0038 (GDPval-AA) at low |
| Knowledge-work drafting at high effort | Either; test both | Within 17 Elo at matched spend on GDPval-AA |
See Claude Haiku 5.5 in Claude Code for subagents and GPT-6 Luna for high-volume API workloads for high-QPS pipelines. Coming from Haiku 4.5? Haiku 5.5 vs Haiku 4.5 lists the five new 400 errors.
Test both side by side in Apidog
The head-to-head that settles it runs on your prompts. In Apidog, create one project with two requests and the same prompt:
- Store
ANTHROPIC_API_KEYand your OpenAI key as environment variables. - Add the Haiku 5.5 request below, then the Luna request with the same prompt and effort.
- Assert that Haiku’s
stop_reasonisn’t"refusal"(there’s no server-side fallback) and thatusagestays under your token budget. - Run both at low, medium and high effort and save each response to compare tokens, latency and output.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"max_tokens": 2048,
"thinking": {"type": "adaptive"},
"output_config": {"effort": "medium"},
"messages": [
{"role": "user", "content": "Classify this support ticket as billing, bug, or feature request: The export button returns a 500 error since this morning."}
]
}'
Leave out temperature, top_p and top_k: non-default values return a 400 on Haiku 5.5. Walkthrough: how to use the Claude Haiku 5.5 API. For eval sets, see testing LLM applications.
FAQ
Are Claude Haiku 5.5 and GPT-6 Luna the same price? Yes, for Haiku prompts up to 100K tokens: $0.10/$0.50, $0.01 cache reads, $0.125 cache writes. Above that, Haiku 5.5 costs $0.50/$2.50 while Luna stays at base until 272K.
Which model scores higher? Haiku 5.5, on every shared row of Anthropic’s launch table. Anthropic compiled that table; no independent head-to-head exists yet.
Which is cheaper per task? Usually Luna: at every effort level on Anthropic’s GDPval-AA chart and four of five on OSWorld (at medium, Haiku 5.5 is $0.0004 cheaper). At matched spend, Haiku 5.5 leads on OSWorld; GDPval-AA is close.
Is either one free? In chat only: Free Claude.ai users can select Haiku 5.5, and Free and Go users get Luna in the ChatGPT desktop app. Neither gives you an API key. See Claude Haiku 5.5 for free.
Pick by prompt length, then test
Check your prompt lengths first. Under 100K tokens, price is a tie and accuracy decides, where Anthropic’s numbers favor Haiku 5.5. Between 100K and 272K, Luna’s threshold saves money on every call. Then run both on real requests at two or three effort levels: download Apidog, put both requests in one project, and let the usage numbers decide.



