Every agent loop you run resends the same opening. The system prompt, the tool definitions, the OpenAPI spec, the repo conventions. The model has already seen that prefix a hundred times today, and you paid full input price for it every single time.
At scale that bill gets loud. OpenAI says its own median researcher spends over $600 a day on coding agents, and the 90th percentile spends $7,000 a day. Those are people running one lab’s internal tools, not a production API serving customers.
Alongside GPT-6 Sol and GPT-6 Luna on September 22, 2026, OpenAI shipped a prompt caching release that goes after exactly this. Cached input reads are discounted 90%, hit rates are higher by default, two of the most common cache-busting changes no longer bust the cache, you can place explicit breakpoints, and there is now a Prompt Caching Dashboard plus a diagnostics tool to see what is actually happening. GitHub, running this in production, reports more than 50% fewer prompt tokens needing fresh processing across billions of requests.
This guide covers what changed, what the discount is worth in real money, how to lay out a prompt so the cache actually hits, and how to test that it keeps hitting after someone refactors your system prompt.
What shipped
| Change | What it means for your code |
|---|---|
| 90% discount on cached input reads | The repeated prefix bills at a tenth of the input rate |
| Higher hit rates by default | Existing integrations get some of this with no code change |
| Changing reasoning effort no longer breaks the cache | You can escalate from low to max mid-session and keep the prefix warm |
| Changing tool availability no longer breaks the cache | Add or remove a tool without paying for a cold prefix |
| Explicit breakpoints | You decide where the cached prefix ends instead of guessing |
| Prompt Caching Dashboard and a diagnostics tool | Hit rate becomes something you measure, not something you assume |
The first three matter most for agents. The last one matters most for the engineer who has to explain last month’s invoice.
The arithmetic
Prompt caching only touches input tokens. Output pricing does not move, so an agent that writes long answers off a short prompt will barely notice. An agent that reads a large stable context and emits a short verdict will notice a lot.
Here is what the discount does to the published rates for the two new models:
| Model | Input per 1M | Cached read per 1M | Output per 1M | Context |
|---|---|---|---|---|
GPT-6 Sol (gpt-6-sol) |
$2.00 | $0.20 | $10.00 | 872k |
GPT-6 Luna (gpt-6-luna) |
$0.10 | $0.01 | $0.50 | 1M |
The cached-read column is the 90% discount applied to the published input rate. Confirm it against the live pricing page before you build a budget on it.
Now a concrete workload. Say you run a code-review agent: 40,000 tokens of stable prefix (system prompt, tool schemas, the service’s OpenAPI document, your API conventions) plus about 2,000 tokens of diff per run, 500 runs a day.
| Sol, no cache hits | Sol, prefix cached | Luna, no cache hits | Luna, prefix cached | |
|---|---|---|---|---|
| Prefix tokens per day | 20M at $2.00 = $40.00 | 20M at $0.20 = $4.00 | 20M at $0.10 = $2.00 | 20M at $0.01 = $0.20 |
| Fresh tokens per day | 1M at $2.00 = $2.00 | 1M at $2.00 = $2.00 | 1M at $0.10 = $0.10 | 1M at $0.10 = $0.10 |
| Daily input cost | $42.00 | $6.00 | $2.10 | $0.30 |
That is roughly 86% off the input bill on Sol for a workload that did not change shape at all. The only thing that changed is whether the prefix stayed byte-identical between calls.
Worth putting next to it: Claude Opus 5.5 lists cached reads at $0.20 per million against $4.00 input, with cache writes at $5.00 per million. The read price lands in the same place as Sol’s, and Anthropic charges an explicit premium to write the cache. OpenAI’s launch does not state a cache write rate for Sol or Luna, which is one of the open questions at the end of this piece. If you are comparing whole bills rather than single numbers, our September 2026 model price war breakdown puts all three launches in one table.
Why cache misses were usually your own fault
Prompt caching matches on a prefix. The cache is keyed on the leading run of tokens in your request, so the first byte that differs from the previous call invalidates everything after it. One character in the wrong place costs you the whole 40,000-token discount. The classic offenders, in rough order of how often they show up in real code:
- A timestamp, request ID or trace ID injected into the system prompt.
- The user’s name, org ID or locale placed above the static instructions instead of below them.
- Retrieved context inserted before the tool definitions, so every new retrieval shifts the whole prefix.
- A tool array built from a set or dictionary, so the JSON key order shuffles between processes.
- Few-shot examples rotated per request for variety.
Two more used to be on that list and are now off it. Changing reasoning effort no longer breaks the cache, so an agent that runs low on the easy turns and escalates to max on the hard one keeps its warm prefix across the escalation. Changing tool availability no longer breaks it either, so gating a dangerous tool behind a permission check no longer forces a cold read of everything above it. Those two patterns are extremely common in agent frameworks, and until now they quietly zeroed out the cache on exactly the turns that cost the most.
Lay the prompt out stable first, volatile last
The discipline survives every future pricing change: sort your context by how often it changes, most stable at the top.
- System or developer instructions that are identical for every user.
- Tool and function schemas, serialized in a fixed order.
- Long shared reference material: the OpenAPI document, the style guide, the database schema.
- Per-tenant or per-project context that changes rarely.
- Retrieved chunks, which change per query.
- The user turn and anything with a clock in it.
Layers one through three are your cache. If a timestamp has to be in the prompt, put it in the user turn, not the system message. If your tool list is assembled dynamically, sort it and freeze the JSON key order so two processes serialize it identically.
from openai import OpenAI
client = OpenAI()
STATIC_PREFIX = [
{"role": "developer", "content": SYSTEM_INSTRUCTIONS}, # never changes
{"role": "developer", "content": OPENAPI_DOCUMENT}, # changes on deploy
]
response = client.responses.create(
model="gpt-6-sol",
reasoning={"effort": "low"},
tools=SORTED_TOOL_SCHEMAS,
prompt_cache_key="review-agent:v7",
input=STATIC_PREFIX + [
{"role": "user", "content": f"Review this diff:\n{diff}"}, # volatile, last
],
)
print(response.usage)
The explicit breakpoints in this release let you mark where that cached prefix ends rather than relying on the platform to infer it. That is most useful when layer four is large: a per-tenant block that is stable for one customer and different for the next, where you want the boundary drawn deliberately instead of at whatever token the heuristic picked.
Version the cache key itself, as in review-agent:v7 above. When you change the prefix on purpose, bump the version, and you will see the cold period in your metrics instead of wondering why costs jumped.
Verify it, do not assume it
The usage object on every response tells you how many input tokens were served from cache. Read it, log it, and alert on it. A rough shape:
usage = response.usage
cached = usage.input_tokens_details.cached_tokens
total_in = usage.input_tokens
print(f"cache hit rate: {cached / total_in:.0%} billed fresh: {total_in - cached}")
Field names in the snippets above come from the current OpenAI API shapes rather than from the launch announcement, so check them against the reference before you ship. The principle does not change: the response tells you the truth, so put that number on a dashboard.
The new Prompt Caching Dashboard and diagnostics tool cover the fleet-level view, which is where you will spot the pattern GitHub reports: more than 50% of prompt tokens no longer needing fresh processing across billions of requests. That is a platform-wide measurement, and the only way to know your own figure is to measure it.
The failure mode is silent. Nothing errors when the cache stops hitting. A teammate adds a request ID to the system prompt for debugging, the prefix goes cold, and the only signal is a slow line on the invoice three weeks later.
That makes it a regression test, not a monitoring problem. Treat the model endpoint like any other API you depend on: save the request in Apidog, assert that usage.input_tokens_details.cached_tokens exceeds a floor on the second call in a scenario, and run it on a schedule in CI. Two calls in sequence, one to warm and one to assert, is the whole test, and a prompt refactor that kills the cache then fails on the pull request instead of on the bill.
The latency footnote
Caching cuts the work done before the first token, so it usually helps time to first token as well as cost. Do not expect it to rescue the reasoning modes.
Artificial Analysis, a third party, measures time to first token at 102.15 seconds for GPT-6 Sol and 124.23 seconds for GPT-6 Luna on their “max” reasoning variants. Those are not the numbers you will see at low effort, and they are not vendor figures, so treat them as directional. The point stands either way: no caching discount turns a two-minute thinking budget into an interactive one. Measure latency per effort level before you promise a response time. Our GPT-6 Astra API guide covers the effort levels in more detail.
What the launch does not answer
Four things are worth confirming before you commit a design to this:
- Whether cache writes carry a charge on Sol and Luna, and at what rate.
- How long a cached prefix lives without traffic, and whether the retention controls from the GPT-5.6 era carry over unchanged.
- Whether GPT-6 Astra’s cached read rate moved to the same 90% discount, since the launch material covers the new models.
- The exact parameter name for explicit breakpoints, which the announcement describes by capability rather than by field.
None of them block the layout work. Stable-first ordering, a versioned cache key and an assertion on cached tokens pay off under any of the answers.
The short version
Sort your prompt by volatility and keep the clock out of the system message. Version the cache key so deliberate changes are visible. Read cached_tokens on every response and put a floor under it in CI. Escalate reasoning effort and toggle tools freely now, because neither costs you the prefix any more. For background on the mechanism itself, our explainer on how prompt caching works covers the fundamentals.
The pricing on these models will move again. It always does. Prompt layout discipline is the part that keeps paying after the next launch resets every number in the table.



