Claude Opus 5.5 Prompt Caching: The Break-Even Math on $0.20 Cached Reads

Claude Opus 5.5 charges $0.20 per million cached input tokens against $4.00 fresh and $5.00 to write. Here is the break-even reuse rate, 21%, worked against real request shapes.

INEZA Felin-Michel

INEZA Felin-Michel

23 September 2026

Claude Opus 5.5 Prompt Caching: The Break-Even Math on $0.20 Cached Reads

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Your agent resends the same 40,000 tokens of system prompt, tool definitions and policy documents on every single call. At Claude Opus 5.5 input pricing of $4 per million tokens, that prefix costs $0.16 every time it moves, whether or not a single byte of it changed since the last request.

Prompt caching is the fix, and Anthropic priced it aggressively on Opus 5.5. A cached read costs $0.20 per million tokens against $4.00 for fresh input, a 95% discount. But a cache write costs $5.00 per million, which is more than sending the tokens uncached. So caching is not free money. It is a bet that you will reuse the prefix, and the bet has a precise break-even point that almost nobody calculates before turning it on.

This piece works that arithmetic out. The short answer: on a single prefix you break even after 1.26 calls, and in steady state you break even at a cache hit rate of about 21%. Below that, caching costs you money.

button

The scale is worth stating first. OpenAI disclosed at its own launch that its median researcher spends over $600 per day on coding agents, with the 90th percentile above $7,000 per day. At that volume a 20 percentage point difference in hit rate is a salary. For the wider pricing picture across all three September launches, see our September 2026 model price war breakdown.

The three rates that decide everything

Token type Claude Opus 5.5 rate per million Relative to fresh input
Fresh input $4.00 baseline
Cache write $5.00 $1.00 premium
Cached read $0.20 $3.80 saving
Output $20.00 caching does not touch it

Two facts fall straight out of that table. Writing a prefix into the cache costs you $1.00 per million more than not caching it at all. Every later read of that prefix saves you $3.80 per million. Output pricing never moves, so an agent that emits long answers from a short prompt has little to gain here, while one that reads a large stable context and returns a short verdict has a great deal to gain.

The full spec sheet, including the 1,000,000 token context window and the 128,000 token maximum output, is in what is Claude Opus 5.5.

The break-even is 1.26 calls, not two

Take a prefix of exactly one million tokens and N calls that share it, one write and N minus one reads.

uncached:  N * $4.00
cached:    $5.00 + (N - 1) * $0.20

4.00N = 5.00 + 0.20(N - 1)
3.80N = 4.80
N     = 1.26

You need 1.26 calls to repay the write premium. Since calls come in whole numbers, that means: if the prefix is ever read back even once, caching it was correct. Two calls already put you 35% ahead.

Calls sharing one write Uncached cost per 1M prefix Cached cost Saving
1 $4.00 $5.00 25% worse
2 $8.00 $5.20 35%
5 $20.00 $5.80 71%
10 $40.00 $6.80 83%
100 $400.00 $24.80 94%

That table assumes one write and perfect reuse afterwards. Real systems are messier, which is where the second calculation comes in.

In steady state, the only number that matters is hit rate

Over a day of traffic you are not writing once. Caches expire, prefixes get edited, new tenants arrive. Let h be the fraction of your prefix tokens served from cache. A miss bills as a write, a hit bills as a read:

effective cost per 1M prefix tokens = $5.00 * (1 - h) + $0.20 * h
                                    = $5.00 - $4.80h

break-even against $4.00 uncached:  h = 1.00 / 4.80 = 20.8%

Twenty one percent is the number to remember. If fewer than roughly one in five of your prefix tokens are served from cache, switching caching on made your bill worse.

Cache hit rate Effective cost per 1M prefix Versus $4.00 uncached
0% $5.00 25% worse
20.8% $4.00 break-even
50% $2.60 35% cheaper
75% $1.40 65% cheaper
90% $0.68 83% cheaper
95% $0.44 89% cheaper
99% $0.25 94% cheaper
100% $0.20 95% cheaper

The two tables are the same equation seen from different ends, since one write per N calls is a hit rate of (N-1)/N. Ten calls per write is a 90% hit rate and both tables say 83%.

Three request shapes, worked

Shape 1: high-frequency agent, small prefix

A support triage agent with a 40,000 token prefix, an 800 token variable user turn, 600 tokens of output, 10,000 calls a day. Assume writes on 3% of calls, so a 97% hit rate.

Uncached Cached
Prefix, 400M tokens/day $1,600.00 $137.60
Variable turn, 8M tokens/day $32.00 $32.00
Input total per day $1,632.00 $169.60

That is 89.6% off the input line, about $1,462 a day, or $43,800 a month. Output stays at $120 a day either way. Note the prefix is only 40,000 tokens. Caching pays here because of frequency, not size.

Shape 2: one-shot pass over a 1M context

One million tokens in, one answer out, never reused. Uncached that call costs $4.00 of input. Cached it costs $5.00, because you paid to write a prefix nobody read. Run 500 documents a day through that pattern and caching costs you an extra $500 a day for nothing.

This is the shape people get wrong most often, because the context is enormous and the instinct is that enormous contexts obviously need caching. Size is irrelevant. Reuse is the whole variable.

Shape 3: long agent session on a 1M window

Now take the same million token context and have an agent re-read it across 200 turns, which is exactly what an 18 hour task loop looks like.

Cost
Uncached, 200 x $4.00 $800.00
Cached, 1 write + 199 reads $44.80
Cached, 5 writes + 195 reads $64.00

Even if the cache goes cold four times mid-session and you pay five full writes, you are still 92% below the uncached bill. Long sessions are where the $0.20 rate earns its reputation.

The expensive mistake: caching the part that changes

Consider 5,000 documents of 60,000 tokens each, summarised once apiece behind a shared 6,000 token instruction header.

Strategy Input cost
No caching at all $1,320
Cache the whole request, including each document $1,650
Cache only the 6,000 token header $1,206

Caching the wrong boundary is 25% worse than not caching. Caching the right one is 9% better. Same feature, same rates, a $444 spread decided entirely by where the cached prefix ends.

The rule that falls out of this: cache the longest leading run of bytes that is identical across calls, and not one byte more. If a timestamp, a request ID or a per-request document sits inside your cached prefix, your hit rate collapses to zero and every call bills at $5.00 instead of $4.00. The general mechanics of prefix matching are covered in our prompt caching primer.

95% is an asymptote, not a discount you receive

The headline number is that cached reads cost 5% of fresh input. You will never actually pay 5%, because you always paid for at least one write. At 100 calls per write you are at 94%. At 10 calls per write you are at 83%. At 2 calls you are at 35%.

Budget from the hit rate table, not from the headline. A finance team that models a 95% saving and observes 83% will conclude the feature is broken when it is working exactly as priced.

What the launch material does not tell you

Three inputs to this arithmetic are not in Anthropic’s Opus 5.5 launch material, and guessing at them would be the fastest way to get a budget wrong:

Check all three on the vendor pricing page before committing to a number. The rates in this article are published. These three are not.

Test that the cache is actually hitting

The failure mode is silent. Nothing throws an error when a prefix goes cold. Someone adds a debug ID to the system message, the hit rate drops from 97% to 0, and the only signal is a line on an invoice three weeks later.

The Messages API reports the split on every response:

"usage": {
  "input_tokens": 812,
  "cache_creation_input_tokens": 0,
  "cache_read_input_tokens": 40960,
  "output_tokens": 604
}

On the first call cache_creation_input_tokens carries the prefix. On every call after it, that field should be 0 and cache_read_input_tokens should carry the same load. Save the request once in Apidog, fire it twice, and watch which field moves.

Then turn that observation into an assertion so it cannot regress quietly. An Apidog test scenario that sends the request twice and asserts cache_read_input_tokens > 40000 on the second response will fail in CI the moment a teammate makes the prefix non-deterministic. That is a one-line test standing between you and a 25x increase in prefix billing, and it is the single highest-return thing in this article.

Where GPT-6 lands on the same math

GPT-6 Sol lists input at $2.00 per million with a 90% discount on cached reads, which puts a cached read at $0.20 per million, the same figure as Opus 5.5. GPT-6 Luna at $0.10 per million input works out to $0.01 per million cached.

The difference is what we can compute. OpenAI’s launch material does not state a cache write rate, so the break-even reuse rate for GPT-6 cannot be derived the way it can for Opus 5.5. Anthropic publishing an explicit $5.00 write price is what makes the 21% figure calculable at all. The rest of OpenAI’s caching release, including cache-preserving effort changes and the new diagnostics tooling, is in our GPT-6 prompt caching write-up.

The bottom line

Prompt caching on Claude Opus 5.5 is one arithmetic question: will more than a fifth of your prefix tokens come back warm? If yes, switch it on and expect 83% to 94% off the input line rather than the advertised 95%. If your workload is one shot passes over unique documents, leave it off and save the $1.00 per million write premium. And whichever you choose, assert on cache_read_input_tokens in CI, because a cache regression never announces itself.

Explore more

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, no shared benchmark. Specs, vendor claims, and how to test both yourself.

30 September 2026

Is GPT-6.1 Sol Free?

Is GPT-6.1 Sol Free?

Is GPT-6.1 Sol free? No: it's for Plus and up in ChatGPT Work and Codex, and the API has no free tier. Here are the cheapest routes, with real cost math.

30 September 2026

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

Is Claude Sonnet 5.5 free? Yes on Claude.ai (web, iOS, Android). Every free route checked, plus what isn't: Claude Code, the API, and Copilot Free.

29 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Claude Opus 5.5 Prompt Caching: The Break-Even Math on $0.20 Cached Reads