Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens on the Claude API, the same as Sonnet 5. Cache reads cost $0.20 per million, the Batch API halves both rates to $1 and $5, the full 1M context window bills at the standard rate, and fast mode isn’t offered. Sonnet 5’s planned September 1 rise to $3/$15 was cancelled, so $2/$10 is now the standard price.
This guide lists every line item of Claude Sonnet 5.5 pricing, works through three cost examples with the arithmetic shown, and explains why effort moves your bill more than the per-token rate. New to the model? Start with what Claude Sonnet 5.5 is, or read how to use Claude Sonnet 5.5 for free. To check the numbers against your own traffic, send the request from Apidog and read the usage block on each response.
Claude Sonnet 5.5 API pricing table
Prices per million tokens (MTok), from Anthropic’s pricing docs:
| Line item | Price per MTok |
|---|---|
| Input | $2.00 |
| 5-minute cache write | $2.50 |
| 1-hour cache write | $4.00 |
| Cache read | $0.20 |
| Output | $10.00 |
| Batch input | $1.00 |
| Batch output | $5.00 |
Three details change what you pay:
- US-only inference.
inference_geo: "us"costs 1.1x on every token category on the Claude API and Claude Platform on AWS ($2.20 input, $11 output, $0.22 cache reads). Foundry’s US Data Zone deployment carries the same 1.1x; Bedrock and Google Cloud price separately. - No long-context premium. A 900k-token request costs the same per token as a 9k-token one.
- Same tokenizer as Sonnet 5. Your token counts carry over unchanged.
How Sonnet 5.5 compares on list price
| Model | Input | Output | Notes |
|---|---|---|---|
| Claude Sonnet 5.5 | $2 | $10 | Cache read $0.20; 1M context; no fast mode |
| Claude Sonnet 5 | $2 | $10 | Cache read $0.20; $3/$15 rise cancelled |
| Claude Opus 5.5 | $4 | $20 | Cache read $0.20; fast mode $8/$40 |
| Claude Haiku 4.5 | $1 | $5 | 200k context |
| GPT-6 Sol | $2 | $10 | 90% off cached input; 872k context |
| GPT-6 Luna | $0.10 | $0.50 | 1M context |
Sonnet 5.5 costs half of Opus 5.5 on input, output and cache writes. It shares the exact $2/$10 list price with GPT-6 Sol, so choosing between those two comes down to results per task. For the backstory, see Claude Sonnet 5 pricing, the September 2026 price war and GPT-6 Sol pricing.
Worked examples: what real requests cost
The formula for every line: tokens ÷ 1,000,000 × price per MTok. Price each line item, then add.
Example 1: one chat-style request
3,000 input tokens, 800 output tokens, no caching.
- Input: 3,000 ÷ 1,000,000 × $2 = $0.006
- Output: 800 ÷ 1,000,000 × $10 = $0.008
- Total: $0.006 + $0.008 = $0.014
A thousand of those cost $14. Output is 21% of the tokens but 57% of the bill, so output length matters more than prompt length.
Example 2: an agent re-reading a 50,000-token prefix
An agent loop sends the same 50,000-token prefix (system prompt, tools, repo context) on 20 calls. Only the prefix is priced here.
Uncached: 20 × 50,000 = 1,000,000 tokens; 1,000,000 ÷ 1,000,000 × $2 = $2.00
Cached, with a 5-minute write:
- Call 1 writes the cache: 50,000 ÷ 1,000,000 × $2.50 = $0.125
- Calls 2 to 20 read it: 19 × 50,000 = 950,000 tokens; 950,000 ÷ 1,000,000 × $0.20 = $0.19
- Total: $0.125 + $0.19 = $0.315
That saves $1.685, about 84%. The 5-minute write pays off only while calls land inside the window; for slower loops, a 1-hour write costs 50,000 ÷ 1,000,000 × $4 = $0.20, so the total is $0.20 + $0.19 = $0.39, still about 80% below uncached. The same reasoning is worked for Opus in our prompt caching cost math; the $0.20 read rate is identical.
Example 3: a batch job of 10,000 requests
Same shape as Example 1, sent through the Batch API.
- Input: 10,000 × 3,000 = 30,000,000 tokens; 30 × $1 = $30
- Output: 10,000 × 800 = 8,000,000 tokens; 8 × $5 = $40
- Batch total: $70, against 30 × $2 + 8 × $10 = $140 at standard rates
Batch also raises the output ceiling from 128K to 300K tokens with the output-300k-2026-03-24 beta.
Cost per task, not per token
The rate card tells you what a token costs, not how many tokens your task burns. That second number swings far more with effort than any rate gap between models.
Anthropic says that in its testing Sonnet 5.5 “costs up to 30% less per task than its predecessor.” The launch post also plots cost at every effort level. The Terminal-Bench 4.0 runs are Anthropic’s own, FrontierCode was run by Cognition, and the index costs come from Artificial Analysis:
| Effort | Terminal-Bench 4.0 (score, $/attempt) | FrontierCode v1.1 (score, $/task) | AA index (score, cost to run) |
|---|---|---|---|
| low | 20.0%, $0.76 | 29.3%, $0.19 | 35.8, $544 |
| medium | 28.8%, $0.83 | 36.5%, $0.24 | 40.7, $701 |
| high | 43.0%, $1.94 | 49.4%, $0.42 | 46.7, $1,176 |
| xhigh | 61.5%, $5.30 | 52.1%, $1.59 | 51.9, $2,738 |
| max | 70.6%, $12.54 | 46.2%, $20.78 | 56.0, $8,977 |
What stands out:
- Max costs about 6.5x high on Terminal-Bench ($12.54 ÷ $1.94 = 6.46). Artificial Analysis’s index cost about 16.5x more to run at max than at low ($8,977 ÷ $544).
- More effort isn’t always better. FrontierCode drops from 52.1% at xhigh to 46.2% at max while cost per task rises 13x ($20.78 ÷ $1.59). Anthropic’s footnote says that at max the model more often ran a review skill that fans out subagents, which caused a timeout or out-of-scope edits in cases Cognition examined.
- Opus at lower effort competes. Opus 5.5 at medium scored 57.6% on Terminal-Bench for $2.94, against Sonnet 5.5’s 43.0% at high for $1.94. Our Sonnet 5.5 vs Opus 5.5 comparison works through that trade.
Verbosity at max is the other trap. OfficeChai’s write-up of Artificial Analysis data puts Sonnet 5.5 at about 193,000 output tokens per index task at max, with cost per task about 50% above Sonnet 5. Simon Willison reported on Hacker News that a max-effort run burned 128,000 thinking tokens and ran out before producing the answer.
Six levers that cut the bill
- Set effort on purpose. The API defaults to
high; Claude Code and the Claude apps default tomedium. Anthropic’s prompting guide suggestsmediumorlowfor chat andxhighormaxonly for measured quality gains. Don’t carry Sonnet 5 settings over; the scale was recalibrated. - Use
between_toolswhere you had thinking off.thinking: {"type": "disabled"}now returns a 400;{"type": "between_tools"}replaces it atlow,mediumandhigh. - Cache anything reused. The minimum cacheable prompt is 512 tokens (1,024 on Sonnet 5). Changing top-level
effortbetween requests invalidates the cache, so hold it steady. - Batch what can wait. Half price on input and output.
- Keep
max_tokenshonest. Anthropic’s guide treatsstop_reason: "max_tokens"as a failed response, so a truncated answer is spend with nothing to ship. - Trim unrequested work. Sonnet 5.5 adds tests, docs and files you didn’t ask for, and starts extra review rounds at
xhighandmax. Anthropic’s suggested system-prompt paragraph cut session cost by about a third atmaxwith no quality change.
Where else you pay for Sonnet 5.5
- Claude plans (claude.com/pricing): Free is $0 and includes Sonnet in chat, with no Claude Code. Pro is $17/month billed annually or $20 monthly. Max starts at $100/month. Team seats are $20/$25 (standard) or $100/$125 (premium), annual/monthly. None of them includes API access.
- GitHub Copilot: Pro ($10/month), Pro+, Max, Business and Enterprise, billed at provider list pricing. Copilot Pro includes 1,500 AI credits a month at $0.01 each, about $15 of tokens, or roughly 1,070 Example 1 requests ($15 ÷ $0.014). Copilot Free and Student can’t select it.
- Cursor: same rates from the Other Models pool (Pro, Pro Plus and Ultra), plus 10% for US-only endpoints ($2.20/$11).
- OpenRouter: $2/$10, or $1/$5 on the
:batchvariant. No:freevariant. - Vercel AI Gateway: $2/$10 with $0.20 cache reads. The $5 free credit doesn’t cover Sonnet 5.5.
- Bedrock, Google Cloud, Foundry: each cloud’s own price page. Bedrock supports the Standard tier only, with no Batch.
Hunting for credits instead? See is there a free Claude Sonnet 5.5 API.
Track cost per request in Apidog
Every Messages response carries a usage object with input_tokens, output_tokens and cache counts. That’s enough to price any call in Apidog:

- Store your key as an
ANTHROPIC_API_KEYenvironment variable. - Save this request once per effort level (
low,medium,high) with the same prompt:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"max_tokens": 4000,
"thinking": {"type": "between_tools"},
"output_config": {"effort": "medium"},
"messages": [{"role": "user", "content": "Classify this support ticket: ..."}]
}'
- Add a post-processor that prices the call and fails on truncation:
const body = pm.response.json();
const u = body.usage;
const cost = (u.input_tokens * 2 + u.output_tokens * 10) / 1e6;
pm.environment.set("last_call_usd", cost.toFixed(5));
pm.test("not truncated", () => pm.expect(body.stop_reason).to.not.eql("max_tokens"));
- Run all three and compare
usage.output_tokensand cost side by side. If you cache, add the cache counts at their own rates.
The full request walkthrough is in how to use the Claude Sonnet 5.5 API.
FAQ
How much does Claude Sonnet 5.5 cost per million tokens? $2 input and $10 output on the Claude API. Cache reads cost $0.20; the Batch API charges $1 and $5.
Is Sonnet 5.5 more expensive than Sonnet 5? No. The per-token price is identical, and Anthropic says Sonnet 5.5 costs up to 30% less per task in its testing.
Is there a surcharge above 200k tokens? No. The full 1M window bills at the standard rate.
Does Sonnet 5.5 have fast mode? No. Fast mode is available only on Opus 5.5, Opus 5 and Opus 4.8.
Is Claude Sonnet 5.5 free? In the Claude chat app, yes: anyone can chat with it on the Free plan. The API runs on prepaid credits; the free API guide covers the credit programs.
Next step: price your own workload
Take your three most common request shapes, run each at two effort levels, and multiply the usage numbers by the table above. You’ll see quickly whether medium holds your quality bar or high earns its extra tokens. Download Apidog to keep those requests and the cost script side by side.



