If you keep a spreadsheet of model prices, it went stale this week. xAI shipped Grok 4.7 on September 21. The next day, Anthropic shipped Claude Opus 5.5 and OpenAI shipped two new GPT-6 models, Sol and Luna. Three launches in 48 hours, two frontier labs on the same day, and every row in your cost model is now wrong.
The per-token numbers are below. The more useful story is that per-token price stopped being the unit anyone is competing on, and that one of the week’s loudest claims expired within hours of publication.
Every new price in one table
| Model | API id | Input $/M | Output $/M | Cached input | Context | AA Intelligence Index |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | claude-opus-5-5 |
$4 | $20 | $0.20 read, $5 write | 1M | 58, #1 of 212 |
| GPT-6 Sol | gpt-6-sol |
$2 | $10 | 90% off reads | 872k | 48 |
| GPT-6 Luna | gpt-6-luna |
$0.10 | $0.50 | 90% off reads | 1M | 37 |
| GPT-6 Astra | n/a | $10 | $50 | n/a | n/a | n/a |
| Claude Opus 5 (previous) | claude-opus-5 |
$5 | $25 | n/a | n/a | n/a |
| Grok 4.7 (Sep 21) | n/a | $2 | $6 | n/a | n/a | n/a |
Three things stand out. Opus 5.5 lands at $4/$20 against Opus 5’s $5/$25, so the flagship got 20% cheaper while taking the top spot on the Artificial Analysis index at 58. Sol sits at a fifth of Astra’s rate on both input and output. And Luna puts a 1M context window at ten cents per million input tokens.
Anthropic also lists a fast mode for Opus 5.5 at $8/$40 and a 128k maximum output. If you have rates hardcoded from our Claude Opus 5 pricing or GPT-5.6 pricing write-ups, those pages now describe the previous generation.
What actually moved: cost per task, not cost per token
OpenAI’s launch post does not lead with per-token rates. It leads with cost per completed task, and the table it publishes is the most interesting artifact of the week.
| Model (reasoning effort) | AutomationBench 1.0.6 | Cost per task |
|---|---|---|
| GPT-6 Sol (xhigh) | 33.2% | $0.27 |
| Fable 5.1 with Opus 5 fallback (max) | 31.4% | more than 8.9x Sol |
| GPT-6 Astra (low) | 30.3% | 3.9x Sol |
| Claude Opus 5 (max) | 26.9% | 11.1x Sol |
Read the bottom row carefully. Sol at extra-high reasoning scores higher than Opus 5 at max for about 9% of Opus 5’s cost per task. Cost per task is the honest unit because it folds in what per-token pricing hides: reasoning tokens, tool calls, retries, and how many turns a model needs before it stops. A model at twice the sticker rate that finishes in a third of the tokens is cheaper, and no pricing page tells you that.
The pattern repeats. On Agents’ Last Exam V1, Sol at max scores 56.4%, above Opus 5’s best result, at 60% lower cost per task. On DeepSWE 1.1, Sol at max hits 68.8%, within 1.1 points of Fable 5 at xhigh (69.9%) and roughly 80% cheaper per task; Luna at max scores 66.6%, which OpenAI puts in the same band as Opus 5 and Fable 5 at medium effort, at 93% less per task than Opus 5 and 96% less than Fable 5. On OSWorld 2.0 offline, Sol at xhigh scores 60.5% against Opus 5 at medium on 60.3%, again around 80% cheaper per task. On factuality, OpenAI says Sol makes about half the mistakes of its predecessor, and Luna at higher effort matches GPT-5.6 Sol at about a hundredth of the cost.
The practical consequence: the procurement question is no longer “what is the rate per million tokens.” It is “what does one of my tasks cost end to end,” and only your own workload can answer it.
The comparison that expired the same day it shipped
Every claim above is measured against Claude Opus 5, because Opus 5.5 did not exist when OpenAI’s post was written. Anthropic shipped Opus 5.5 hours later: 20% cheaper than Opus 5, 40% less to run by Anthropic’s own figure, 30% faster on output, and top of the Artificial Analysis index.
So “beats Opus 5 at 9% of the cost per task” is a true statement about a model superseded the same afternoon. Nobody has published a Sol-versus-Opus-5.5 cost-per-task figure, because every harness in these tables ran before Opus 5.5 existed. If you are about to move a workload on the strength of that ratio, you are comparing against a model you would no longer deploy.
Be careful in the other direction too. Anthropic’s own sheet for Opus 5.5 reads: Terminal-Bench 4.0 66.4%, FrontierCode v1.1 54.4%, CursorBench 4.0 57.8%, GDPval-AA v2.1 at 1846 Elo, AutomationBench 40.0%, Humanity’s Last Exam 67.7% with tools, OSWorld 2.0 81.8%, Chartography 89.0%. That AutomationBench 40.0% is not subtractable from OpenAI’s 33.2%: different harness version, different effort settings, different run, no stated common baseline. Two vendors quoting the same benchmark name is not a head-to-head. OpenAI’s own position is that Astra “continues to be our best model across the board,” so the GPT-6 story this week is about price and efficiency, not a new ceiling. Comparisons circulating on X that put specific Opus 5.5 numbers against specific Astra numbers are tweets, not vendor data.
“50% cheaper” is measured against a discount, and the real cut is bigger
OpenAI states that Sol and Luna are each 50% cheaper than GPT-5.6 promotional pricing. Promotional is OpenAI’s word, and it is load bearing: the baseline being halved is a discounted rate, not a standard list rate. A lot of this week’s coverage dropped the qualifier.
Here is the lineage, using the GPT-5.6 list prices we documented in our own launch coverage at the time:
| Tier | GPT-5.6 list (in / out per 1M) | GPT-5.6 promotional | GPT-6 |
|---|---|---|---|
| Sol | $5 / $30 | $4 / $20 | $2 / $10 |
| Terra | $2.50 / $15 | $2 / $12 | no GPT-6 Terra announced |
| Luna | $1 / $6 | $0.20 / $1.20 | $0.10 / $0.50 |
Two things fall out. First, the honest version of the headline beats the headline: against GPT-5.6 list pricing, Sol at $2/$10 is a 60% cut on input and 67% on output, and Luna at $0.10/$0.50 is 90% and 92%. OpenAI compared against its own promo rate, which understates the move for anyone budgeting against list.
Second, there is no GPT-6 Terra. GPT-5.6 was Sol, Terra and Luna; GPT-6 is Astra, Sol and Luna. If you have a middle tier pinned in config, no like-for-like successor is waiting for it, and our GPT-5.6 naming guide now reads as history rather than a current map.
The cheap models are not the fast models
Artificial Analysis measured the new models. For the max reasoning variants, it reports GPT-6 Sol at 115.2 output tokens per second with a time to first token of 102.15 seconds, and Luna at 153.9 tokens per second with 124.23 seconds to first token. [VERIFY] These are third-party figures, not vendor figures, and they describe the highest reasoning setting, not what you would see at low or medium effort.
Directionally, this is the trade the price war is offering: frontier-adjacent quality at a fraction of the cost, paid for in latency before the first byte arrives. Fine for a batch job or a nightly agent run. Fatal for a request with a human waiting.
Measure this yourself rather than inherit it from a leaderboard, because it moves with your effort setting, prompt size and region. Point Apidog at each candidate model’s completions endpoint, send the same prompt at the effort level you would actually ship, and record time to first byte and total response time next to the body. Saved as a test scenario, the next launch is a re-run rather than a research project.
Caching is the other half of this week’s price cut
Alongside the models, OpenAI shipped a real prompt caching release for GPT-6: a 90% discount on cached input reads, higher hit rates by default, a Prompt Caching Dashboard, a diagnostics tool, and two behavioral fixes that matter more than the discount. Changing reasoning effort or tool availability no longer invalidates the cache, and explicit breakpoints let you decide where a cached prefix ends. OpenAI cites GitHub seeing more than 50% fewer prompt tokens needing fresh processing across billions of requests. Anthropic’s side is in the table above: Opus 5.5 cache reads cost $0.20 per million against $4 for fresh input, so a cached read is 5% of a fresh one.
If your agent re-sends a large system prompt and a fat tool schema on every turn, and most do, these changes will move your bill more than any sticker price here. That makes caching the most durable item in the week and the one worth engineering against.
What to do this week
- Move model ids into config, not code. Three ids changed in two days.
claude-opus-5-5,gpt-6-solandgpt-6-lunabelong in environment variables you can flip, not string literals in a client wrapper. - Re-run your own eval before switching. Vendor cost-per-task tables are measured on vendor harnesses against vendor tasks. Yours will differ.
- Instrument cost per task. Log input, output and reasoning tokens against completed units of work, retries included. Per-token dashboards will mislead you for the rest of this year.
- Set a latency budget and test at shipping effort. A model that is 93% cheaper per task and two minutes slower to first token is a different product, not a cheaper one.
- Audit your cache boundary. Stable prefix first, volatile content last, explicit breakpoints where the platform supports them.
- Update contract tests for the new envelopes. Sol is 872k context, Luna and Opus 5.5 are 1M, and Opus 5.5 caps output at 128k. Assertions written against last month’s limits pass silently and truncate in production.
- Route by task instead of migrating wholesale. The spread between $0.10 and $10 per million input tokens is two orders of magnitude. Cheap tiers for extraction and classification, expensive tiers for work that needs them.
One number for scale: OpenAI disclosed that its median internal researcher spends over $600 per day on coding agents, with the 90th percentile at $7,000. Cuts of this shape are aimed at workloads with that headroom. If yours is smaller, they mostly buy you the option to run a better model on the same budget.
Availability, briefly
Sol and Luna are in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu. Free and Go get Luna in the desktop app. Neither is in Chat yet. The API ids are gpt-6-sol and gpt-6-luna, alongside the existing GPT-6 Astra API. Opus 5.5 runs on the Claude Platform, AWS, Google Cloud and Azure.
For model-by-model detail: what is Claude Opus 5.5, what is GPT-6 Sol, what is GPT-6 Luna, the Sol versus Opus 5.5 head-to-head, and the GPT-6 prompt caching release.
The bottom line
Sticker prices fell, but the real event is that the industry switched units. Cost per task, cache hit rate and time to first token now decide what a model costs you, and none of them appear on a pricing page. Measure all three on your own traffic this week, because the comparison tables everyone is quoting were out of date by the time they were published.



