Gemini 4 Argon leads 13 of the 19 rows in Google’s launch table outright, ties 1 (CWE-bench v1, with GPT-6 Astra), and trails on 5: FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0. The catch is access. Argon is available today only to Fairwind Program defenders, so nobody outside Google’s partner list and a few benchmark orgs can rerun these numbers yet.
Below: the full table with who measured each row, the five losses, the fine print, the cyber scores, third-party scoreboards, and how to build your own eval in Apidog before access opens. For the model overview, start with what Gemini 4 Argon is. For the buying decision, see Argon vs GPT-6 Astra vs Claude Opus 5.5.
The full table, with who measured each row
These are Google-reported results from the launch post and the evals methodology PDF, which is headed “results as of October, 2026.” Google computed 10 of the 19 Argon scores itself; the other 9 come from public leaderboards. Rival numbers come from those leaderboards, Google’s own runs, or the vendors’ system cards and blog posts, and harnesses differ by row. Bold marks the row leader.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 | Who measured it |
|---|---|---|---|---|---|
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% | Vals AI |
| AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% | Zapier leaderboard (private set) |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% | Vals AI |
| Harvey Legal Agent | 19.6% | 5.4% | 6.7% | 3.8% | Vals AI |
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% | Argon self-computed; Astra leaderboard; Anthropic system cards |
| FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% | Proximal leaderboard |
| Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% | Vals AI |
| Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% | Argon self-computed; rivals leaderboard |
| PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% | Self-computed for all models |
| Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% | Argon self-computed; rivals leaderboard |
| LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% | Self-computed for all models |
| RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% | Surge leaderboard |
| GraphWalks up to 128K | 99.7% | 98.7% | 91.4% | 90.6% | Self-computed for all models |
| GraphWalks 256K to 1M | 84.2% | 71.8% | 65.0% | 66.8% | Self-computed for all models |
| Agent’s Last Exam | 39.5% | 34.2% | n/r | 38.2% | Argon self-computed; rivals leaderboard |
| OSWorld-2.0 (offline) | 69.2% | 72.6% | n/r | n/r | Argon self-computed; Astra from OpenAI’s blog post |
| Chartography | 71.6% | 71.0% | 46.2% | 66.3% | Surge leaderboard |
| LVBench | 91.7% | 87.5% | 79.7% | 83.7% | Self-computed for all models |
| CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% | Public leaderboard |
n/r = not reported. Grok 4.7, which isn’t in Google’s table, also scores 68% on the CWE-bench leaderboard, so that tie is three-way.
Sorted by source, 9 rows come from public leaderboards for every model (Vals AI, Zapier, Proximal, Surge and CWE-bench). Argon leads 7 of those 9 and ties 1. Google ran 5 rows itself for all four models, and Argon leads 4. The other 5 pair a Google-run Argon score with rival numbers from elsewhere, and that’s where Argon loses most: 3 of its 5 losses sit in those mixed rows. If self-computing flattered Argon, you’d expect the opposite.
Where Argon trails, and why it matters for agentic coding
- FrontierSWE v2: 55.0% against Astra’s 65.5%. Argon is last of four, behind Fable 5.1 (56.3%) and Opus 5.5 (62.3%), on a leaderboard that scored every model.
- Terminal-bench 4.0: 57.4% against Opus 5.5’s 66.4%. Last again, though Astra and Fable 5.1 are within a point.
- PostTrainBench: 45.3% against Opus 5.5’s 49.3%. Second place, on a row Google ran for all models.
- Terminal-Bench Science 0.1: 57.6% against Astra’s 68.1%. Third, ahead of Fable 5.1 only. Google ran Argon’s attempt with a 6x verifier timeout.
- OSWorld-2.0 (offline subset): 69.2% against Astra’s 72.6%. Anthropic reports only a combined online and offline score, so neither Claude model appears.
Agentic coding splits 2 to 2. Argon wins DeepSWE v1.1 and Vibe Code Bench, then finishes last on FrontierSWE v2 and Terminal-bench 4.0. The DeepSWE win mixes harnesses: Argon ran on a mini-swe agent harness, Astra’s number is from the public leaderboard, and the Claude numbers are from system cards. Vibe Code Bench is cleaner, since Vals AI scored all four, but the margin is 1.6 points.
If your workload is terminal-heavy agents or long repository tasks, weigh those two last-place rows above the headline count. Astra holds the FrontierSWE lead; our GPT-6 Astra hands-on shows what that model is like to use today. For Anthropic’s side, the Claude Fable 5.1 benchmarks breakdown covers Fable’s own launch numbers.
Bloomberg reports that anonymous Google employees with access say Argon underwhelms on some coding and front-end design work relative to its benchmark scores. Google called that characterization inaccurate.
Methodology caveats that change how you read the table
- Max effort on both sides. Argon ran “with the Gemini API with the highest thinking settings.” For Astra, Fable 5.1 and Opus 5.5, Google defaults to the “maximum thinking/reasoning settings available.” That compares ceilings, not default behavior or cost.
- LVBench frames. Google ran LVBench itself for all four models, without tools. Gemini sampled video at 1 FPS. Astra got 800 frames, Fable 5.1 got 300 and Opus 5.5 got 600, “due to API limitations.” The models didn’t see identical inputs.
- OSWorld best of three. Argon’s score is “maxed over 3 runs with a single attempt per run,” a best run rather than an average.
- Agent’s Last Exam window. Argon ran on the ALE-Claw harness with a 5-hour window.
- Safety filters on. Agent’s Last Exam and OSWorld ran with safety filters enabled, and flagged responses came back as empty strings. If anything, that works against Argon.
The cyber rows from DeepMind’s cyber page
The DeepMind cyber page adds four scores beyond the launch table:
| Cyber eval | Gemini 4 Argon | Comparison |
|---|---|---|
| Real-world vulnerability discovery | 85.8% | Gemini 3.8 Flash Cyber 71.0% |
| Wiz Penetration Test Benchmark | 70.9% | Gemini 3.8 Flash Cyber 58.2% |
| Gray Swan IPI attack success (k=15, lower is better) | 0.7% | Lowest on the chart; Kimi K3 highest at 52.7% |
| CWE-bench v1 | 68% | Three-way tie with Grok 4.7 and GPT-6 Astra; Opus 5.5 at 67% |
The vulnerability eval used an internal Antigravity harness “that is not cyber specialized,” with source access. The Wiz test gives the model “access only to the running website and its public behavior, without application source code.” Both headline comparisons are against Google’s previous cyber model, not rivals. Our Argon cyber defense explainer covers what these mean for the APIs you run.
Third-party scoreboards
These orgs posted scores within minutes of launch, so they had pre-release access.
- Artificial Analysis lists Gemini 4 Argon (High) at an Intelligence Index of 53, #8 of 223. That rank counts each reasoning setting separately. By distinct model, Argon ties Fable 5.1 and GPT-6 Astra and sits behind Opus 5.5 (58 at max) and Sonnet 5.5 (56 at max). AA reports a 15% hallucination rate on AA-Omniscience (Astra at max, 51%), with accuracy of 50% against 63%. Running the index cost $1.99 per task, at about 62K output tokens per task versus about 27K for Astra at max. Artificial Analysis shows no speed or latency data.
- Vals AI puts Argon at 68.90% on the Vals Index, #1 of 41 and the first Gemini model to top it, at $15.68 per test. Vals AI lists a 262K max output for the configuration it tested, below Google’s stated 1M.
- Arena ranks Argon #1 on Text at 1525, marked Preliminary on 4,942 votes, and #8 on WebDev.
Why outlets publish 12, 13 and 14 wins
Outlets have printed three different win counts. Here’s the recount from Google’s 19 rows:
| Counted against | Argon wins | Ties | Argon loses | Not reported |
|---|---|---|---|---|
| Best of all three rivals | 13 | 1 | 5 | 0 |
| GPT-6 Astra only | 14 | 1 | 4 | 0 |
| Claude Opus 5.5 only | 14 | 0 | 4 | 1 |
| Claude Fable 5.1 only | 15 | 0 | 2 | 2 |
A 13 is right against the whole field. A 14 is right head to head against Astra or Opus 5.5. A 12 matches none of these framings of the full 19 rows, so check which rows that source counted. Margins vary too: Harvey Legal Agent (19.6% against 6.7%) is wide; Chartography is 0.6 points.
How to run your own eval the day access opens
Benchmarks describe Google’s harness; your prompts describe your product. Build the eval today on a model you can call, then point it at Argon with one change.
- Pick real prompts that match the rows you care about: a repository fix, a terminal task, a long-document question.
- In Apidog, create an environment with
GEMINI_API_KEYandGEMINI_MODELset togemini-3.8-flash. Google hasn’t published Argon’s model ID, so this variable is the one value you’ll change. - Save each prompt as a request that reads the model from
{{GEMINI_MODEL}}, grouped into a test scenario. - Add assertions on the output (status, required fields, expected content) and on
usageMetadata, including a ceiling onthoughtsTokenCountsized to your cost cap. - Run the scenario on 3.8 Flash now and keep the report as your baseline.
The day Argon’s ID ships, swap the variable and rerun. Google says “all new models” will launch on the Interactions API, so save an Interactions version of each request too. Our Argon vs Gemini 3.8 Flash comparison covers what to expect on price and limits.
FAQ
Are Gemini 4 Argon’s benchmarks independently verified? Partly. Nine of the 19 rows come from public leaderboards for every model, and Artificial Analysis, Vals AI and Arena posted their own results. The rest are Google-run for Argon.
What is Gemini 4 Argon’s DeepSWE score? 77.9% on DeepSWE v1.1, ahead of Opus 5.5 (74.2%) and Astra (74.1%). Argon’s run used a mini-swe agent harness, while rival numbers come from a leaderboard and system cards.
Does Argon beat GPT-6 Astra on benchmarks? Head to head in Google’s table, Argon wins 14 rows, ties 1 and loses 4. On Artificial Analysis they tie at 53.
Can I rerun these benchmarks myself? Not yet. Argon is limited to Fairwind partners, and Google hasn’t given a public release date; our Argon release date tracker follows the rollout.
Next step
Build the scenario now, run it on 3.8 Flash, and keep the report. When Argon opens up, you’ll have your own numbers minutes after the swap. Download Apidog to set it up.



