Gemini 4 Argon Benchmarks: All 19 Rows, How Google Ran Them, and the 5 It Loses

Gemini 4 Argon benchmarks: all 19 rows with who measured each, the 5 it loses, Google's methodology caveats, and AA, Vals and Arena scores.

INEZA Felin-Michel

INEZA Felin-Michel

2 October 2026

Gemini 4 Argon Benchmarks: All 19 Rows, How Google Ran Them, and the 5 It Loses

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Gemini 4 Argon leads 13 of the 19 rows in Google’s launch table outright, ties 1 (CWE-bench v1, with GPT-6 Astra), and trails on 5: FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0. The catch is access. Argon is available today only to Fairwind Program defenders, so nobody outside Google’s partner list and a few benchmark orgs can rerun these numbers yet.

Below: the full table with who measured each row, the five losses, the fine print, the cyber scores, third-party scoreboards, and how to build your own eval in Apidog before access opens. For the model overview, start with what Gemini 4 Argon is. For the buying decision, see Argon vs GPT-6 Astra vs Claude Opus 5.5.

The full table, with who measured each row

These are Google-reported results from the launch post and the evals methodology PDF, which is headed “results as of October, 2026.” Google computed 10 of the 19 Argon scores itself; the other 9 come from public leaderboards. Rival numbers come from those leaderboards, Google’s own runs, or the vendors’ system cards and blog posts, and harnesses differ by row. Bold marks the row leader.

Benchmark Gemini 4 Argon GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5 Who measured it
Vals Index 68.9% 63.1% 65.8% 67.0% Vals AI
AutomationBench 51.3% 41.4% 31.4% 42.5% Zapier leaderboard (private set)
Vals Finance Agent v2 65.4% 53.5% 58.9% 58.6% Vals AI
Harvey Legal Agent 19.6% 5.4% 6.7% 3.8% Vals AI
DeepSWE v1.1 77.9% 74.1% 67.4% 74.2% Argon self-computed; Astra leaderboard; Anthropic system cards
FrontierSWE v2 55.0% 65.5% 56.3% 62.3% Proximal leaderboard
Vibe Code Bench 91.9% 89.6% 90.3% 90.3% Vals AI
Terminal-bench 4.0 57.4% 58.2% 57.9% 66.4% Argon self-computed; rivals leaderboard
PostTrainBench 45.3% 44.3% 40.2% 49.3% Self-computed for all models
Terminal-Bench Science 0.1 57.6% 68.1% 52.6% 63.3% Argon self-computed; rivals leaderboard
LABBench 2 88.8% 85.4% 68.6% 73.1% Self-computed for all models
RiemannBench 76.0% 72.0% 65.6% 69.6% Surge leaderboard
GraphWalks up to 128K 99.7% 98.7% 91.4% 90.6% Self-computed for all models
GraphWalks 256K to 1M 84.2% 71.8% 65.0% 66.8% Self-computed for all models
Agent’s Last Exam 39.5% 34.2% n/r 38.2% Argon self-computed; rivals leaderboard
OSWorld-2.0 (offline) 69.2% 72.6% n/r n/r Argon self-computed; Astra from OpenAI’s blog post
Chartography 71.6% 71.0% 46.2% 66.3% Surge leaderboard
LVBench 91.7% 87.5% 79.7% 83.7% Self-computed for all models
CWE-bench v1 68.0% 68.0% 58.0% 67.0% Public leaderboard

n/r = not reported. Grok 4.7, which isn’t in Google’s table, also scores 68% on the CWE-bench leaderboard, so that tie is three-way.

Sorted by source, 9 rows come from public leaderboards for every model (Vals AI, Zapier, Proximal, Surge and CWE-bench). Argon leads 7 of those 9 and ties 1. Google ran 5 rows itself for all four models, and Argon leads 4. The other 5 pair a Google-run Argon score with rival numbers from elsewhere, and that’s where Argon loses most: 3 of its 5 losses sit in those mixed rows. If self-computing flattered Argon, you’d expect the opposite.

Where Argon trails, and why it matters for agentic coding

Agentic coding splits 2 to 2. Argon wins DeepSWE v1.1 and Vibe Code Bench, then finishes last on FrontierSWE v2 and Terminal-bench 4.0. The DeepSWE win mixes harnesses: Argon ran on a mini-swe agent harness, Astra’s number is from the public leaderboard, and the Claude numbers are from system cards. Vibe Code Bench is cleaner, since Vals AI scored all four, but the margin is 1.6 points.

If your workload is terminal-heavy agents or long repository tasks, weigh those two last-place rows above the headline count. Astra holds the FrontierSWE lead; our GPT-6 Astra hands-on shows what that model is like to use today. For Anthropic’s side, the Claude Fable 5.1 benchmarks breakdown covers Fable’s own launch numbers.

Bloomberg reports that anonymous Google employees with access say Argon underwhelms on some coding and front-end design work relative to its benchmark scores. Google called that characterization inaccurate.

Methodology caveats that change how you read the table

The cyber rows from DeepMind’s cyber page

The DeepMind cyber page adds four scores beyond the launch table:

Cyber eval Gemini 4 Argon Comparison
Real-world vulnerability discovery 85.8% Gemini 3.8 Flash Cyber 71.0%
Wiz Penetration Test Benchmark 70.9% Gemini 3.8 Flash Cyber 58.2%
Gray Swan IPI attack success (k=15, lower is better) 0.7% Lowest on the chart; Kimi K3 highest at 52.7%
CWE-bench v1 68% Three-way tie with Grok 4.7 and GPT-6 Astra; Opus 5.5 at 67%

The vulnerability eval used an internal Antigravity harness “that is not cyber specialized,” with source access. The Wiz test gives the model “access only to the running website and its public behavior, without application source code.” Both headline comparisons are against Google’s previous cyber model, not rivals. Our Argon cyber defense explainer covers what these mean for the APIs you run.

Third-party scoreboards

These orgs posted scores within minutes of launch, so they had pre-release access.

Why outlets publish 12, 13 and 14 wins

Outlets have printed three different win counts. Here’s the recount from Google’s 19 rows:

Counted against Argon wins Ties Argon loses Not reported
Best of all three rivals 13 1 5 0
GPT-6 Astra only 14 1 4 0
Claude Opus 5.5 only 14 0 4 1
Claude Fable 5.1 only 15 0 2 2

A 13 is right against the whole field. A 14 is right head to head against Astra or Opus 5.5. A 12 matches none of these framings of the full 19 rows, so check which rows that source counted. Margins vary too: Harvey Legal Agent (19.6% against 6.7%) is wide; Chartography is 0.6 points.

How to run your own eval the day access opens

Benchmarks describe Google’s harness; your prompts describe your product. Build the eval today on a model you can call, then point it at Argon with one change.

  1. Pick real prompts that match the rows you care about: a repository fix, a terminal task, a long-document question.
  2. In Apidog, create an environment with GEMINI_API_KEY and GEMINI_MODEL set to gemini-3.8-flash. Google hasn’t published Argon’s model ID, so this variable is the one value you’ll change.
  3. Save each prompt as a request that reads the model from {{GEMINI_MODEL}}, grouped into a test scenario.
  4. Add assertions on the output (status, required fields, expected content) and on usageMetadata, including a ceiling on thoughtsTokenCount sized to your cost cap.
  5. Run the scenario on 3.8 Flash now and keep the report as your baseline.

The day Argon’s ID ships, swap the variable and rerun. Google says “all new models” will launch on the Interactions API, so save an Interactions version of each request too. Our Argon vs Gemini 3.8 Flash comparison covers what to expect on price and limits.

FAQ

Are Gemini 4 Argon’s benchmarks independently verified? Partly. Nine of the 19 rows come from public leaderboards for every model, and Artificial Analysis, Vals AI and Arena posted their own results. The rest are Google-run for Argon.

What is Gemini 4 Argon’s DeepSWE score? 77.9% on DeepSWE v1.1, ahead of Opus 5.5 (74.2%) and Astra (74.1%). Argon’s run used a mini-swe agent harness, while rival numbers come from a leaderboard and system cards.

Does Argon beat GPT-6 Astra on benchmarks? Head to head in Google’s table, Argon wins 14 rows, ties 1 and loses 4. On Artificial Analysis they tie at 53.

Can I rerun these benchmarks myself? Not yet. Argon is limited to Fairwind partners, and Google hasn’t given a public release date; our Argon release date tracker follows the rollout.

Next step

Build the scenario now, run it on 3.8 Flash, and keep the report. When Argon opens up, you’ll have your own numbers minutes after the swap. Download Apidog to set it up.

button

Explore more

Gemini 4 Argon vs Gemini 3.8 Flash: Should You Wait for Argon or Ship on Flash Now?

Gemini 4 Argon vs Gemini 3.8 Flash: Should You Wait for Argon or Ship on Flash Now?

Gemini 4 Argon vs Gemini 3.8 Flash: price, output limits, shared benchmarks and the cost of one call. Ship on Flash now or wait for Argon?

2 October 2026

Gemini 4 Argon's 1M Output Tokens: What a Million-Token Response Does to Your API Stack

Gemini 4 Argon's 1M Output Tokens: What a Million-Token Response Does to Your API Stack

Gemini 4 Argon's 1M output tokens is an output limit, not a context window. What a maxed response costs, plus streaming, timeouts, and caps.

2 October 2026

Gemini 4 Argon API: What's Confirmed, What It Will Cost, and How to Get Your Code Ready

Gemini 4 Argon API: What's Confirmed, What It Will Cost, and How to Get Your Code Ready

Gemini 4 Argon API: no public access or model ID yet. What Google confirmed on price and output, and code to get ready on Gemini 3.8 Flash today.

2 October 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Gemini 4 Argon Benchmarks: All 19 Rows, How Google Ran Them, and the 5 It Loses