GPT-6.1 Sol and Claude Sonnet 5.5 both list at $2 per million input tokens and $10 per million output tokens, and they launched a day apart: Sonnet 5.5 on September 28, 2026, and GPT-6.1 Sol on September 29. Neither vendor has benchmarked one against the other. OpenAI compared GPT-6.1 Sol with Claude Opus 5.5 and Fable 5.1, and Anthropic compared Sonnet 5.5 with GPT-6 Sol, the previous Sol. The honest way to choose is to run both on your own tasks and compare cost per task, not price per token.
Below: specs and prices side by side, what each vendor claims and against which model, the third-party numbers (all on GPT-6 Sol, not 6.1), and a method to run the comparison yourself in Apidog. For each model on its own, see what is GPT-6.1 Sol and what is Claude Sonnet 5.5.
GPT-6.1 Sol vs Claude Sonnet 5.5: specs and pricing
Sources: OpenAI’s pricing page, Anthropic’s Sonnet 5.5 overview and pricing, plus the GPT-6.1 Sol model page linked below.
| GPT-6.1 Sol | Claude Sonnet 5.5 | |
|---|---|---|
| Model ID | gpt-6.1-sol |
claude-sonnet-5-5 (Bedrock: anthropic.claude-sonnet-5-5) |
| Released | Sep 29, 2026 | Sep 28, 2026 |
| Input / output per 1M | $2 / $10 | $2 / $10 |
| Cache reads per 1M | $0.10 | $0.20 |
| Cache writes per 1M | $2.50 | $2.50 (5-minute), $4 (1-hour) |
| Batch per 1M | $1 / $5 | $1 / $5 |
| Prompts over 272K input tokens | $4 input, $0.20 cached, $15 output for the whole request | No premium across the 1M window |
| Faster tier | Fast at $4 / $20; Ultrafast “coming soon” | Fast mode not available |
| Context window | 1,050,000 (922,000 max input) | 1M |
| Max output | 128,000 | 128K (300K on Batch with a beta header) |
| Knowledge cutoff | Apr 30, 2026 | June 2026 |
| Effort levels | low, medium (default), high, xhigh, max; no none |
low, medium, high (API default), xhigh, max; thinking can’t be turned off (lowest setting: between_tools) |
| API shape | Responses (with tools), Chat Completions (no tools), Batch | Messages, Batch |
| Other platforms | OpenRouter (openai/gpt-6.1-sol) |
Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| Free access | None. Plus, Pro, Business, Enterprise and Edu get it in ChatGPT Work and Codex (not Chat; Enterprise and Edu once an admin enables it, per the models docs); no free API tier | Anyone can chat with it on Claude.ai; the API bills per token |
Two rows matter most for cost. GPT-6.1 Sol’s cache reads cost half of Sonnet 5.5’s, so reusing a long system prompt is cheaper on OpenAI. Above 272K input tokens the balance flips: the GPT-6.1 Sol model page prices the full request at 2x input and cache rates and 1.5x output, while Sonnet 5.5 stays at $2 and $10. A 400,000-token prompt with a 5,000-token answer costs about $1.68 on GPT-6.1 Sol and $0.85 on Sonnet 5.5, uncached.
What each vendor claims, and against which model
| OpenAI on GPT-6.1 Sol | Anthropic on Claude Sonnet 5.5 | |
|---|---|---|
| Compared against | GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1 | Sonnet 5, Opus 5.5, GPT-6 Sol (GPT-5.6 Sol in two charts) |
| Includes the other model | No | No (GPT-6.1 Sol didn’t exist yet) |
| Headline | “Near-Astra intelligence” at one-fifth of Astra’s standard token prices | 30%+ faster than Sonnet 5, up to 30% less per task |
| Who ran the numbers | OpenAI (competitor results from public reports, per its footnote) | Anthropic, plus Cognition, Cursor, Artificial Analysis and Surge AI for named benchmarks |
OpenAI’s GPT-6.1 Sol launch post states these OpenAI-reported results in its text:
- AutomationBench 1.0.6, medium effort: +2.2 pp over Opus 5.5 at roughly a third of the cost, and +4.8 pp over GPT-6 Sol.
- Terminal-Bench Science 0.1, max effort: $5.47 per task, versus $23.21 for Opus 5.5. GPT-6 Astra still scores highest, at 68.1%.
- OSWorld 2.0 offline, max effort: +7 pp over GPT-6 Sol for less than half the cost.
Anthropic’s Sonnet 5.5 launch post includes a GPT-6 Sol column:
- FrontierCode 1.1, run by Cognition: Sonnet 5.5 scores 52.1% at xhigh (46.2% at max) versus 49.3% for GPT-6 Sol. At high effort, Anthropic says Sonnet 5.5 matches GPT-6 Sol’s best score for about one-fifth the cost.
- GDPval-AA v2.1 and AA-Briefcase v1.1, run by Artificial Analysis: 1844 versus 1487, and 1811 versus 1483.
The Sonnet 5.5 benchmarks breakdown covers the rest of Anthropic’s table.
Opus 5.5 appears on both charts, and both vendors report AutomationBench and Terminal-Bench Science, so it’s tempting to bridge the two. The numbers don’t line up. OpenAI reports AutomationBench 1.0.6 at medium effort in its own harness, while Anthropic’s system card summary table runs its Claude models at max effort (Sonnet 5.5 44.7, Opus 5.5 42.5; GPT-6 Sol 32.0). Anthropic gives Sonnet 5.5 59.9% on Terminal-Bench Science, but OpenAI’s text gives no raw score for GPT-6.1 Sol. And OpenAI tested computer use on OSWorld 2.0 offline, Anthropic on OSWorld 2.1. The GPT-6 Sol vs Claude Opus 5.5 comparison hit the same wall.
What third parties measured (on GPT-6 Sol, not 6.1)
Every third-party comparison in search results pairs Sonnet 5.5 with GPT-6 Sol. The most complete is Artificial Analysis, whose Intelligence Index v4.3.2 aggregates 10 evaluations:
| Effort | Sonnet 5.5 index | Sonnet 5.5 cost per task | GPT-6 Sol index | GPT-6 Sol cost per task |
|---|---|---|---|---|
| medium | 41 | $0.59 | 40 | $0.25 |
| high | 47 | $1.08 | 43 | $0.38 |
| xhigh | 52 | $2.74 | 44 | $0.52 |
| max | 56 | $7.60 | 48 | $1.05 |
At max effort, the same page measures Sonnet 5.5 at 138 output tokens per second and GPT-6 Sol at 76. Sonnet 5.5 scores higher and costs more per task at every effort shown, despite identical list prices; token consumption is the difference. Sonnet 5.5 at high (47, $1.08) sits next to GPT-6 Sol at max (48, $1.05).
Anthropic’s own chart data cuts both ways. On AA-Briefcase, GPT-6 Sol costs less per task at every effort ($0.34 versus $1.64 at medium) while Sonnet 5.5 scores higher. On FrontierCode, Sonnet 5.5 at high reaches 49.4% for $0.42 per task, against GPT-6 Sol’s 49.3% at max for $2.07.
All of this is the older Sol. OpenAI says GPT-6.1 Sol beats GPT-6 Sol at the same or lower effort on several benchmarks, so expect its row to move once third parties test it.
Where each is likely stronger, by the vendors’ own framing
GPT-6.1 Sol. OpenAI pitches it for “complex coding, computer use, and professional work when you want near-Astra performance at a lower cost.” Its stated gains are on agentic automation, computer use and science tasks, plus fewer factual errors at low effort. On price it favors cache-heavy workloads under 272K input tokens, and it’s the only one of the two with a paid speed tier (Fast).
Claude Sonnet 5.5. Anthropic pitches it for well-scoped everyday tasks, bug fixes, and polished docs, slides and spreadsheets, and says Opus 5.5 “remains clearly stronger at complex, open-ended work” (see Sonnet 5.5 vs Opus 5.5). Its reported strengths are agentic coding (70.6% on Terminal-Bench 4.0 at max, Anthropic-run), knowledge work, and computer use (80.1% partial on OSWorld 2.1). On price it favors prompts above 272K tokens, and it runs on Bedrock, Google Cloud and Microsoft Foundry.
Two behaviors skew a naive test. Default effort differs (medium for GPT-6.1 Sol, high for Sonnet 5.5 on the API), so compare matching settings. And neither takes sampling knobs: Sonnet 5.5 returns 400 on non-default temperature, top_p or top_k, and OpenAI’s GPT-6 guidance says to drop temperature and top_p when effort isn’t none. Control variance with repeated runs.
Run the comparison yourself in Apidog
Pick 20 to 50 tasks from your real traffic with checkable answers, such as a JSON field or a label. Then in Apidog:
- Create an environment with
OPENAI_API_KEYandANTHROPIC_API_KEYas secrets, plusEFFORT. - Save two requests that share a
{{prompt}}variable. The shapes differ (Responses reference, Messages reference); the API formats comparison walks through why:
POST https://api.openai.com/v1/responses
Authorization: Bearer {{OPENAI_API_KEY}}
Content-Type: application/json
{"model": "gpt-6.1-sol", "reasoning": {"effort": "{{EFFORT}}"},
"max_output_tokens": 25000, "input": "{{prompt}}"}
POST https://api.anthropic.com/v1/messages
x-api-key: {{ANTHROPIC_API_KEY}}
anthropic-version: 2023-06-01
content-type: application/json
{"model": "claude-sonnet-5-5", "max_tokens": 25000,
"output_config": {"effort": "{{EFFORT}}"},
"messages": [{"role": "user", "content": "{{prompt}}"}]}
- Add assertions for each shape, then one on the answer itself against an
expectedcolumn:
| Check | OpenAI Responses | Anthropic Messages |
|---|---|---|
| Finished normally | $.status equals completed |
$.stop_reason equals end_turn |
| Answer present | $.output[*].type contains message |
$.content[*].type contains text |
| Output tokens, reasoning included | $.usage.output_tokens |
$.usage.output_tokens |
| Cache reads | $.usage.input_tokens_details.cached_tokens |
$.usage.cache_read_input_tokens |
| Cache writes | $.usage.input_tokens_details.cache_write_tokens |
$.usage.cache_creation_input_tokens |
- Price each response in a post-processor script. OpenAI’s
input_tokensincludes cached and cache-write tokens, so subtract them before applying the $2 rate (OpenAI prompt caching). Anthropic’sinput_tokenscounts only tokens after the last cache breakpoint (Anthropic prompt caching). Apply the 272K rule on the OpenAI side. - Put both requests in one test scenario, keep prompts on a single line, without double quotes, in a CSV, and run it from the Apidog CLI at each effort you care about:
apidog run --access-token "$APIDOG_ACCESS_TOKEN" -t "$SCENARIO_ID" -e "$ENV_ID" \
-d prompts.csv --env-var "EFFORT=medium" -r cli,junit
Divide total spend by the number of tasks that passed: that’s your cost per task. Run at least medium and high on both, since the same effort name isn’t calibrated the same way across vendors. For request details, see the GPT-6.1 Sol API guide and how to use the Claude Sonnet 5.5 API.
FAQ
Is GPT-6.1 Sol cheaper than Claude Sonnet 5.5? Same list price, $2/$10 per 1M. GPT-6.1 Sol’s cache reads are cheaper; Sonnet 5.5 is cheaper above 272K input tokens. Real cost depends on tokens used per task.
Which is better at coding? There’s no shared benchmark. Anthropic reports Sonnet 5.5 above GPT-6 Sol on FrontierCode; OpenAI reports GPT-6.1 Sol beating GPT-6 Sol’s best DeepSWE score by 6.4 pp. Test on your own repo.
Which is faster? Artificial Analysis measured Sonnet 5.5 at 138 tokens per second and GPT-6 Sol at 76, both at max effort. There’s no third-party number for GPT-6.1 Sol yet.
Is either one free? GPT-6.1 Sol has no free tier. Sonnet 5.5 is available in the Claude.ai chat apps, but the API bills per token. See how to use Claude Sonnet 5.5 for free.
Pick by running both
Take ten tasks from your backlog, run both models at medium and high, and compare cost per passing task. Download Apidog to keep the pair as one scenario you can rerun when Artificial Analysis publishes GPT-6.1 Sol numbers or either vendor ships a new snapshot.



