Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. A prompt of over 100,000 tokens pays higher prices: $0.50 input and $2.50 output. Cache reads start at $0.01, the Batch API halves every rate, and Anthropic says the model costs “around 75% less to run” than Haiku 4.5 on average.
This guide covers every line item, the math behind that 75%, three worked examples (one crossing 100K), cost per effort level, and how Haiku 5.5 compares with Haiku 4.5, GPT-6 Luna and Sonnet 5.5. New to the model? Start with what Claude Haiku 5.5 is. To check the numbers on your own traffic, send requests from Apidog and read each response’s usage block.
Claude Haiku 5.5 API pricing table
Prices per million tokens (MTok), from Anthropic’s pricing docs:
| Line item | Prompts up to 100K tokens | Prompts over 100K tokens |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| 5-minute cache write | $0.125 | $0.625 |
| 1-hour cache write | $0.20 | $1.00 |
| Cache read (hits and refreshes) | $0.01 | $0.05 |
| Batch input | $0.05 | $0.25 |
| Batch output | $0.25 | $1.25 |
The docs put it plainly: “a prompt of over 100,000 tokens pays higher prices.” Other Claude models from 4.6 onward bill the full 1M window at one rate; Haiku 5.5 is the exception. Also note:
- US-only inference. US-only routing via
inference_geocosts 1.1x on every token category on the Claude API and Claude Platform on AWS: $0.11 input and $0.55 output under 100K. - A new tokenizer. The same text produces about 30% more tokens than on Haiku 4.5.
- Smaller tool overhead. The tool-use system prompt is 286 tokens (auto/none), down from 496 on Haiku 4.5.
- No Priority Tier and no fast mode.
Where “around 75% less” comes from
Haiku 5.5 is priced 90% lower than Haiku 4.5 up to 100,000 tokens and 50% lower above; 90% of Haiku 4.5 requests fell under 100K; and the figure nets out the new tokenizer. A rough reconstruction:
- Weight by request count: 0.9 × 90% + 0.1 × 50% = 86% lower.
- Add about 30% more tokens: the remaining 14% of cost becomes 14% × 1.3 = 18.2%, so the saving falls to about 82%.

Anthropic’s figure is lower again. It doesn’t publish the weighting, and long prompts likely carry more of the spend than of the request count. Treat 75% as an average; your usage data is the real answer.
Cost per attempt at each effort level
Haiku 5.5 is the first Haiku with effort levels (low to max; medium is the API default), and effort moves your bill more than the rate card. These figures come from the launch post’s per-effort charts. Anthropic ran OSWorld and Terminal-Bench; Artificial Analysis ran GDPval-AA:
| Effort | OSWorld 2.1 (score, $/attempt) | GDPval-AA v2.1 (Elo, $/task) | Terminal-Bench 4.0 (score, $/attempt) |
|---|---|---|---|
| low | 42.0%, $0.0695 | 1125, $0.0117 | 12.7%, $0.424 |
| medium | 53.3%, $0.1257 | 1277, $0.0304 | 20.3%, $0.6798 |
| high | 61.3%, $0.1827 | 1420, $0.0894 | 24.8%, $1.0372 |
| xhigh | 67.6%, $0.2792 | 1513, $0.2673 | 31.5%, $1.7545 |
| max | 72.4%, $0.6111 | 1620, $0.8659 | 39.2%, $2.6433 |
What stands out:
- Max costs 4.9x medium on OSWorld ($0.6111 ÷ $0.1257) and 28.5x medium on GDPval-AA ($0.86593 ÷ $0.03042).
- Launch-table scores are max effort. At the default
medium, GDPval-AA is 1277, not 1620. - For agentic coding, check Sonnet. On Terminal-Bench, Sonnet 5.5 at
low(with $0.10 cache reads) scored 20% for $0.6205, against Haiku 5.5 atmediumwith 20.3% for $0.6798. Anthropic says Sonnet and Opus “remain better choices for complex agentic coding tasks.” Our benchmarks breakdown has every row.
How Haiku 5.5 compares on price
| Model | Input / output per MTok | Cache read | 5m cache write | Long-prompt rule |
|---|---|---|---|---|
| Claude Haiku 5.5 | $0.10 / $0.50 up to 100K | $0.01 | $0.125 | Over 100K: $0.50 / $2.50 |
| GPT-6 Luna | $0.10 / $0.50 | $0.01 | $0.125 | Over 272K input: 2x input and cache, 1.5x output ($0.20 / $0.75) |
| Claude Haiku 4.5 | $1 / $5 | $0.10 | $1.25 | One rate (200K context) |
| Claude Sonnet 5.5 | $2 / $10 | $0.10 | $2.50 | One rate across 1M |
GPT-6 Luna matches Haiku 5.5 up to 100K; the threshold differs. Per OpenAI’s model page, the Luna surcharge starts above 272K and applies “for the full request.” Example 3’s call stays on Luna’s base rate: (150,000 × $0.10 + 2,000 × $0.50) ÷ 1,000,000 = $0.016, against $0.08 on Haiku 5.5. Tokenizers differ too: one Hacker News commenter estimated 100K Claude tokens at about 60-65K GPT tokens. On Anthropic’s OSWorld chart, Haiku 5.5 at medium (53.3%, $0.1257) beat Luna at max (48.9%, $0.205). More in Claude Haiku 5.5 vs GPT-6 Luna.
Sonnet 5.5 cut cache reads from $0.20 to $0.10 the same day, about 20% cheaper on most agentic work per Anthropic. Our Sonnet 5.5 pricing guide predates the cut.
Haiku 4.5 isn’t deprecated; Haiku 5.5 vs Haiku 4.5 lists the breaking changes to fix first.
API credits and consumer plans
Max and Team API credits. Max 5x gets $100 a month in Claude Platform credits, Max 20x $200, and Team $20 per Standard seat and $100 per Premium seat, pooled and capped at $500. Free, Pro and Enterprise aren’t eligible. Per the API credits docs, they cover the Claude API, not Claude Code or the clouds, and don’t roll over. At $0.10 per MTok, $100 buys 1 billion input tokens, or about 220,000 Example 1 requests ($100 ÷ $0.000455).
Consumer plans (claude.com/pricing): Free ($0), Pro, Max, Team and Enterprise users can select Haiku 5.5 on Claude.ai on web, iOS and Android. Pro is $17/month billed annually or $20 monthly; Max starts at $100/month. Free has no Claude Code, and a chat plan can’t drive a script or CI job. How to use Claude Haiku 5.5 for free covers the rest.
Clouds. Bedrock, Google Cloud and Foundry bill Haiku 5.5 through their own price pages.
Track cost per request in Apidog
Every Messages response returns a usage object, which is enough to price a call in Apidog:
- Store your key as an
ANTHROPIC_API_KEYenvironment variable. - Save this request once per effort level (
low,medium,high) with the same prompt:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"max_tokens": 4000,
"thinking": {"type": "adaptive"},
"output_config": {"effort": "medium"},
"messages": [{"role": "user", "content": "Classify this support ticket: ..."}]
}'
- Add a post-processor that picks the tier and prices the call. It treats
input_tokensas prompt length, which holds without caching:
const body = pm.response.json();
const u = body.usage;
const long = u.input_tokens > 100000;
const rateIn = long ? 0.50 : 0.10;
const rateOut = long ? 2.50 : 0.50;
const cost = (u.input_tokens * rateIn + u.output_tokens * rateOut) / 1e6;
pm.environment.set("haiku_call_usd", cost.toFixed(6));
pm.test("not a refusal", () => pm.expect(body.stop_reason).to.not.eql("refusal"));
- Run all three and compare output tokens and cost side by side.
The refusal check matters because Haiku 5.5’s safety classifiers have no server-side fallback. The full walkthrough is in how to use the Claude Haiku 5.5 API.
FAQ
How much does Claude Haiku 5.5 cost per million tokens? $0.10 input and $0.50 output for prompts up to 100K tokens; $0.50 and $2.50 above.
Is Haiku 5.5 cheaper than Haiku 4.5? Yes: 90% cheaper per token up to 100K, 50% above, and around 75% on average per Anthropic.
Is Haiku 5.5 the same price as GPT-6 Luna? Up to 100K, yes. Above that Haiku 5.5 rises 5x; Luna stays flat until 272K. Our GPT-6 Luna overview has the details.
Do Max plan credits work in Claude Code? No. They cover the Claude Platform API, not Claude Code. See Claude Haiku 5.5 in Claude Code for plan access there.
Next step: price your own workload
Pull your three most common request shapes, check how many cross 100K, and run each at low and medium. Multiply the usage numbers by the table above to get your real saving. Download Apidog to keep those requests and the cost script in one project.



