If you moved an agent workload onto GPT-6 Astra in early September, you have three weeks of invoices by now and you already know the shape of the problem. Astra bills $10 per million input tokens and $50 per million output. A long-running loop with a fat system prompt and a tool schema attached burns through that faster than any spreadsheet predicted.
On September 22, OpenAI shipped GPT-6 Sol at $2 and $10. Same model family, same API surface, one fifth the price on every line of the bill.
This is the migration. What changes in your code, which is close to nothing. What changes in your bill, which is everything. And the part most launch-day coverage skipped: OpenAI still says Astra is the better model, and the one published head to head between the two is not the comparison it looks like.
TL;DR
gpt-6-astratogpt-6-solis a model string swap for most callers. Context window, max output, endpoints, built-in tools, supported features and rate-limit tiers are identical.- Every rate falls by exactly 5x: input $10 to $2, cached input $1 to $0.20, cache writes $12.50 to $2.50, output $50 to $10. Your bill divides by five regardless of token mix.
- Sol adds a
nonereasoning effort. It also restricts Chat Completions function calling toreasoning_effort: "none", the one change that can break a working integration. - Sol’s knowledge cutoff is April 20, 2026. Astra’s is April 30, 2026.
- OpenAI says Astra “continues to be our best model across the board.” The published Sol-versus-Astra benchmark runs Astra at
lowagainst Sol atxhigh, so it measures cost efficiency, not capability ceiling.
The price table
Both sets of rates come from OpenAI’s model pages, gpt-6-astra and gpt-6-sol, read on September 23, 2026.
| Metric, per 1M tokens | GPT-6 Astra | GPT-6 Sol | Change |
|---|---|---|---|
| Input | $10 | $2 | 5x cheaper |
| Cached input | $1 | $0.20 | 5x cheaper |
| Cache writes | $12.50 | $2.50 | 5x cheaper |
| Output | $50 | $10 | 5x cheaper |
The billing modifiers match on both models. Prompts over 272K input tokens bill at 2x the input and cache rates and 1.5x output for the whole request. Batch and Flex are half price. Fast mode is double, and on Astra it carries no latency SLA.

Because all four rates drop by the same factor, you do not need to model your token mix to predict the saving. Take a concrete agent workload: 10,000 requests a day, each with a 30,000-token cached prefix, 10,000 tokens of fresh input and 3,000 tokens of output.
| Component | Daily tokens | Astra | Sol |
|---|---|---|---|
| Cached input | 300M | $300 | $60 |
| Fresh input | 100M | $1,000 | $200 |
| Output | 30M | $1,500 | $300 |
| Total | $2,800 | $560 |
Shift the mix toward output, shift it toward cache, run at 272K context and pay the long-prompt multiplier: the ratio stays at five.
For scale on why that matters, OpenAI reported that its own median researcher spends over $600 a day on coding agents, with the 90th percentile at $7,000 a day. Divide those by five and the number of experiments a team can afford changes. The wider context for both launches is in our September 2026 model price war breakdown.
One qualifier to keep attached: OpenAI describes Sol as 50% cheaper than GPT-5.6, and the comparison is against GPT-5.6 promotional pricing, which is OpenAI’s own word. Against the GPT-5.6 list rates we documented at the time in our GPT-5.6 pricing post, the cut is larger. Against Astra it is a straight 5x.
What stays exactly the same
This is the section that makes the migration cheap.
| GPT-6 Astra | GPT-6 Sol | |
|---|---|---|
| Model ID | gpt-6-astra |
gpt-6-sol |
| Context window | 1,050,000 | 1,050,000 |
| Max input tokens | 922,000 | 922,000 |
| Max output tokens | 128,000 | 128,000 |
| Modalities | text, image in; text out | text, image in; text out |
| Endpoints | Chat Completions, Responses, Batch | Chat Completions, Responses, Batch |
| Not supported | Realtime, Assistants, fine-tuning, embeddings, audio | same |
| Built-in tools | web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, tool search | same list |
| Features | streaming, structured outputs, function calling, file search, image input, web search, prompt caching | same list |
| Tier 5 rate limits | 15,000 RPM, 40M TPM | 15,000 RPM, 40M TPM |
| Snapshots | gpt-6-astra |
gpt-6-sol |
The context window is the headline. Sol is not a shorter-context model: it carries the same 1,050,000-token window and the same 922,000-token input ceiling as Astra. Nothing about your chunking, your retrieval budget or your compaction strategy has to move.
What actually changes in your code
Four things, in the order they are likely to bite.
1. Chat Completions function calling. On Astra, Chat Completions works and tool calling requires the Responses API. On Sol, Chat Completions supports function calling only when reasoning_effort is "none". If you are calling tools over Chat Completions with any other effort, that request stops doing what it did. OpenAI’s own GPT-6 guide says to use Responses for reasoning with tools. If you are already on Responses, this one costs you nothing.
2. The none effort level. Astra supports low through max. Sol supports all of those plus none, which is the lever that makes it viable for classification and extraction work where reasoning tokens are pure overhead. Default on both is medium.
3. Knowledge cutoff. Astra is trained to April 30, 2026, Sol to April 20, 2026. Ten days is small, but if a prompt assumes knowledge from late April, test the assumption.
4. Unsupported parameters. Whenever reasoning effort is not none, temperature, top_p and top_logprobs must be absent, and Chat Completions also drops logprobs. Astra enforces the same rule, so a clean Astra integration already complies. It only matters if you move to reasoning_effort: "none" on Sol and think about putting temperature back.
Here is the before and after for a typical Responses call. The diff is one line.
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
- model="gpt-6-astra",
+ model="gpt-6-sol",
reasoning={"effort": "xhigh"},
tools=[{"type": "function", "name": "run_api_test", "parameters": {...}}],
input=[
{"role": "developer", "content": "You are a senior API engineer. Bias towards action."},
{"role": "user", "content": "Read this OpenAPI operation and propose three negative test cases."},
],
)
Note the effort level in that example. Moving to Sol and keeping the same effort is not the interesting migration. Moving to Sol and turning the effort up is, because you have five times the budget to spend on reasoning tokens for the same money.
What you give up
Be honest about this part, because the launch numbers are easy to misread.
OpenAI says Astra is still the better model. The launch post states that Astra “continues to be our best model across the board.” That is the vendor’s own framing of its own new release, and it is the sentence to quote at anyone who tells you Sol replaces Astra.
The published head to head is not a capability comparison. On AutomationBench 1.0.6, Sol at xhigh scores 33.2% at $0.27 per task, and Astra at low scores 30.3% at 3.9 times Sol’s cost per task. Read the effort levels. Sol is dialed all the way up, Astra is dialed all the way down. What that pairing demonstrates is that Sol’s ceiling clears Astra’s floor at roughly a quarter of the cost per task, which is a real and useful result. It says nothing about Sol at xhigh against Astra at max. No published number covers that matchup. If your workload is one where Astra at high effort was the thing that finally made it work, Sol is a test, not a swap.
Latency at the top end. Artificial Analysis measured GPT-6 Sol’s max-reasoning variant at 115.2 output tokens per second with a 102.15-second time to first token. That figure is third party, not from OpenAI, and it describes the max variant specifically, so it does not tell you what medium or none do. Treat it as a warning that the cheap model is not automatically the fast model at high effort, and measure your own effort level rather than inheriting the number.
Availability. Sol reaches ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, and is not in Chat yet. The API is ready; the Chat surface is not.
For what Astra does that justifies keeping it somewhere in the stack, our two-day hands-on test, the computer-use write-up and the Critical cyber threshold explainer all hold, and the full spec sheet is in our GPT-6 Astra API guide.
Decide with your own requests, not with benchmarks
AutomationBench does not run your prompts. The only comparison that settles a migration is the same request set, sent to both model IDs, scored on your own criteria. Set that up once and it pays for itself on every future launch. In Apidog, put the model ID in an environment variable, save the request once, and switch environments to retarget it:
{
"model": "{{MODEL_ID}}",
"reasoning": { "effort": "xhigh" },
"input": [
{ "role": "user", "content": "{{TEST_PROMPT}}" }
]
}
Build a test scenario from 20 or 30 real production prompts, add assertions for the response shape your parser depends on (output_text present, tool call arguments valid against your JSON Schema, no truncation at max_output_tokens), then run it twice, once per environment. Apidog records the body and the elapsed time for every request, so you get correctness and latency side by side without writing a harness. The usage block on each response gives you the token counts to price the comparison properly.
Two assertions are worth adding for this migration specifically: check that tool calls still arrive if you were on Chat Completions, and judge on your slowest prompt rather than your average one, because the latency risk sits at high effort on long inputs.
Migration checklist
- Confirm you are on the Responses API anywhere you call tools. If you call tools over Chat Completions, move before you switch models.
- Swap
gpt-6-astratogpt-6-soland leave everything else alone for the first run. - Re-run your regression set against both IDs and diff the outputs, not just the status codes.
- Try one step up in reasoning effort on Sol. You have the budget for it now.
- Re-check any prompt that depends on knowledge from late April 2026.
- Keep an Astra path behind a flag for the tasks where the ceiling is what you were paying for.
The bottom line
Astra to Sol is the rare migration where the API surface does not move, the context window does not shrink, and the price falls by a fixed factor on every meter. The work is not in the code. It is in the twenty prompts you run through both models to find out whether your hardest task was using Astra’s headroom or just paying for it.
Run that comparison before you flip the flag, and keep OpenAI’s own sentence in view while you read the results: Astra is still their best model. Sol is the one you can afford to leave running.



