Grok 4.7 is the cheapest of the three on output tokens ($6 per million, against $10 for GPT-6 Sol and $20 for Claude Opus 5.5), and Opus 5.5 leads on every coding benchmark where two of these vendors publish the same name. On the Artificial Analysis Intelligence Index, a third-party composite, the order is Opus 5.5 at 58, Sol at 48, and Grok 4.7 at 46. Which one costs you less depends on caching and on how many tokens each model spends per task; on a cache-heavy agent workload, Sol edges out Grok.
All three shipped within about 36 hours. SpaceXAI released Grok 4.7 on September 21, 2026, and Anthropic and OpenAI released Opus 5.5 and Sol on September 22. That timing created a problem you should know about before reading any comparison, this one included. Below: the spec sheet, the benchmark rows that overlap, the price math with and without caching, routing advice, and a way to test all three in one Apidog project. For each model on its own, see what Grok 4.7 is, what GPT-6 Sol is, and what Claude Opus 5.5 is.
Nobody benchmarked these three against each other
Each launch compares against models that existed when it was written:
- SpaceXAI’s Grok 4.7 post compares against Grok 4.6, GPT-5.6 Sol, and Claude Fable 5.1, plus GPT-6 Astra on two charts. Watch the name: GPT-5.6 Sol is the previous generation, listed at $4/$20 on that page, not the GPT-6 Sol released the next day.
- OpenAI’s Sol launch compares against Claude Opus 5, which Opus 5.5 replaced hours later.
- Anthropic’s Opus 5.5 launch publishes scores with no competitor columns.
So no vendor table has all three in it. What you can do is line up the rows that share a benchmark name, read them as direction rather than a ranking, and use a third-party index as a tiebreaker. Our GPT-6 Sol vs Claude Opus 5.5 breakdown goes deeper on that pair.
The spec sheet
| Grok 4.7 | GPT-6 Sol | Claude Opus 5.5 | |
|---|---|---|---|
| API model id | grok-4.7 |
gpt-6-sol |
claude-opus-5-5 |
| Released | Sep 21, 2026 | Sep 22, 2026 | Sep 22, 2026 |
| Input / output per 1M | $2 / $6 | $2 / $10 | $4 / $20 |
| Cached input read per 1M | $0.50 | $0.20 (90% off) | $0.20 |
| Context window | 500,000 | 872,000 | 1,000,000 |
| Max output | none listed | not stated | 128,000 |
| AA Intelligence Index (third party) | 46 (#21 of 211) | 48 | 58 (#1 of 212) |
| AA output speed (third party) | 82.6 tokens/s at xhigh | 115.2 tokens/s at max, 102 s to first token | not in our sources |
| Where to call it | xAI API, Cursor, GitHub Copilot, Grok Build, OpenRouter, Vercel | OpenAI API, ChatGPT Work, Codex | Claude Platform, AWS, Google Cloud, Azure |
Two rows decide most workloads. Output price: Grok 4.7’s $6 is 40% below Sol and 70% below Opus 5.5, and that matters for reasoning models because thinking tokens bill as output. Context: Grok 4.7 has the smallest window, and xAI bills the whole request at double rates ($4 input, $12 output) once a prompt reaches 200,000 tokens.
The benchmarks that overlap
These rows share a benchmark name across vendor sheets. Each vendor ran its own model on its own harness at its own effort setting, so small gaps are noise. Large gaps still tell you something.
| Benchmark (vendor-run) | Grok 4.7 | GPT-6 Sol | Claude Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 37.6% (xhigh) | not published | 66.4% |
| CursorBench 4.0 | 46.3% (xhigh) | not published | 57.8% |
| DeepSWE v1.1 | 71.0% (high) | 68.8% (max) | not published |
| GDPval, Elo | 1,695 (xhigh) | not published | 1,846 (GDPval-AA v2.1) |
How to read it:
- Terminal work favors Opus 5.5 by a wide margin. A 28.8-point gap on Terminal-Bench 4.0 is too large to explain with harness differences alone. Grok 4.7’s 37.6% is a big jump from Grok 4.6’s 20.3%, and still well behind.
- CursorBench points the same way at 11.5 points.
- DeepSWE is a wash. Grok 4.7 at high effort and Sol at max sit 2.2 points apart, inside what different harnesses produce.
- GDPval versions may differ. Grok’s chart doesn’t state one, so treat the 151-point gap as a hint.
Grok 4.7 does lead on two evals from its own sheet: Harvey’s Legal Agent Benchmark at 19.6% (GPT-5.6 Sol 2.5%, Fable 5.1 6.7%) and EEBench, an electrical engineering eval, at 64.0% (Fable 5.1 56.4%; GPT-6 Astra 69.3% on the same chart). Neither Sol nor Opus 5.5 has a published score on either, so read them as signals for legal and engineering knowledge work, not wins over these two rivals.
One independent harness does run two of the three the same way. On SWE-Together, 109 real-repository tasks run through one agent harness, Opus 5.5 scores 68.8% pass@1 and Grok 4.7 64.7%, about 4 points apart instead of the 29 on Terminal-Bench. GPT-6 Sol isn’t listed there yet. Grok 4.7 also spent $7.81 per task on that board, so its token price didn’t make it cheap per task.
The third-party index agrees with the overall order: Opus 5.5 58, Sol 48, Grok 4.7 46. For every row and caveat on Grok’s side, see our Grok 4.7 benchmarks breakdown, which also covers a benchmark maintainer’s report that Grok 4.7 got around their sandbox to fetch upstream fixes.
What each one costs on a real workload
Per-token prices only tell part of the story. Here’s arithmetic on published rates for two workload shapes at 1,000 runs a day. Cache writes are excluded; Opus 5.5 charges $5 per million for those.
Agent loop, cache-friendly: 50,000 input tokens per run, 45,000 of them a stable prefix (system prompt, tool schemas, docs), plus 3,000 output tokens.
| Daily cost | Grok 4.7 | GPT-6 Sol | Claude Opus 5.5 |
|---|---|---|---|
| No cache hits | $118.00 | $130.00 | $260.00 |
| Prefix cached | $50.50 | $49.00 | $89.00 |
Without caching, Grok 4.7 is cheapest. With the prefix cached, Sol edges ahead, because its cached reads cost $0.20 per million against Grok’s $0.50. The more your prompts repeat, the less Grok’s output discount helps.
Reasoning-heavy task: 10,000 input tokens and 20,000 output tokens per run.
| Daily cost | Grok 4.7 | GPT-6 Sol | Claude Opus 5.5 |
|---|---|---|---|
| Uncached | $140 | $220 | $440 |
Here the output price dominates: Grok 4.7 costs about 64% of Sol and 32% of Opus 5.5.
Both tables assume every model spends the same number of tokens, and they don’t. SpaceXAI’s own CursorBench data shows Grok 4.7 averaging 70,141 output tokens per task at xhigh and 15,677 at low. A model that needs twice the tokens loses its price advantage. Cost per completed task on your prompts decides this; our price war analysis explains why vendors started publishing that metric.
Which one to route to
Start from the workload, then test:
- Grok 4.7 fits output-heavy agent loops, budget-sensitive reasoning, legal and engineering knowledge work, and teams already in Cursor, GitHub Copilot, or Grok Build, where it’s one model-picker change away. Keep prompts under 200,000 tokens.
- Claude Opus 5.5 fits the hardest coding and terminal tasks, where its lead is widest, and anything that needs a 1M window or 128,000-token outputs. It also ran on AWS, Google Cloud, and Azure from day one, which helps if procurement matters.
- GPT-6 Sol fits cache-heavy agents and automation, plus prompts between 200,000 and 872,000 tokens. Budget for the 102-second time to first token Artificial Analysis measured at max effort; our Sol latency guide covers timeouts.
- GPT-6 Luna, at $0.10/$0.50, belongs in the conversation for high-volume extraction and classification, where none of the three above is cost-effective.
Many teams will run two of these behind a router: a cheap default and an expensive escalation path for the tasks that fail.
Test all three in one Apidog project
The comparison that counts is your prompts on your harness. Apidog turns that into a saved test instead of a weekend project:
- Create three environments, one per provider, each holding
base_url, the API key as a secret, andmodel. Grok 4.7 and GPT-6 Sol both use the Responses API (POST {{base_url}}/responseswith aninputfield), so one request definition covers both. Opus 5.5 uses Anthropic’s Messages API (POST /v1/messageswithmessages, a requiredmax_tokens, andx-api-keyplusanthropic-versionheaders), so give it its own request. - Use the same prompt text in every request, pulled from a shared variable so it can’t drift.
- Add assertions for status 200, a non-empty answer, and the presence of
usage.input_tokensandusage.output_tokens. - Run them as one test scenario and compare response time, token counts, and output side by side. Multiply tokens by the price rows above to get cost per run on your task.
Run it at the effort level you’d ship with, not the one on the benchmark chart. Download Apidog to build it.
FAQ
Is Grok 4.7 better than GPT-6? Not on the numbers available. The Artificial Analysis index puts GPT-6 Sol at 48 and Grok 4.7 at 46, and OpenAI calls GPT-6 Astra its best model. Grok 4.7 is cheaper on output and leads on its own legal and engineering evals.
Is Grok 4.7 better than Claude Opus 5.5? Opus 5.5 leads on Terminal-Bench 4.0 (66.4% vs 37.6%) and CursorBench 4.0 (57.8% vs 46.3%) and ranks first on the Artificial Analysis index. Grok 4.7 costs half as much on input and less than a third on output.
Which is cheapest? Per token, Grok 4.7 on output, tied with Sol on input. With heavy prompt caching, Sol can come out ahead because its cached reads cost $0.20 per million against Grok’s $0.50.
Can I try all three for free? Each has limited no-cost routes. See how to use Grok 4.7 for free, GPT-6 Sol for free, and Claude Opus 5.5 for free.
How do I call Grok 4.7 from code? POST https://api.x.ai/v1/responses with "model": "grok-4.7". The Grok 4.7 API guide has curl and SDK examples.
Decide with your own numbers
Pick the two models your workload points to, save the same prompt against both in Apidog, and run it at your production effort level. Compare cost per completed task and latency, then route. When the next model lands, add an environment and re-run.



