Gemini 3.8 Flash is available today: a stable model with a free tier, priced at $0.75 input and $3.75 output per 1M tokens through 2026-12-31, then $1.50 and $7.50, with a 64K output cap. Gemini 4 Argon isn’t. Google’s launch post prices it at $2 and $10 during an intro period of unstated length, then $4 and $20. Google says its output limit is 1M tokens, and it’s available today only to Fairwind Program defenders. If you need to ship this quarter, ship on Flash and keep Argon one config change away.
This post compares the two on specs, the benchmark rows both launch posts share, cyber scores and the cost of an identical call, then gives you a decision rule. For background, see what Gemini 4 Argon is and what Gemini 3.8 Flash is. The last section shows the Apidog setup that makes the switch a single variable.
Side by side: what you can call today
| Gemini 3.8 Flash | Gemini 4 Argon | |
|---|---|---|
| Availability | Stable on the Gemini API, AI Studio, Antigravity and Gemini Enterprise | A subset of Fairwind Program partners, as a managed model on Gemini Enterprise |
| Model ID | gemini-3.8-flash |
Not published |
| Price per 1M tokens (in / out) | $0.75 / $3.75 through 2026-12-31, then $1.50 / $7.50 | $2 / $10 intro (end date not stated), then $4 / $20 |
| Cached input per 1M | $0.075, then $0.15 | $0.10 intro, then $0.20 (95% off input) |
| Output limit | 65,536 tokens | 1M tokens (Google’s stated limit) |
| Input window | 1,048,576 tokens | Not published |
| Thinking levels | low, medium (default), high; minimal returns HTTP 400 |
Not documented |
| Free API tier | Yes, rate-limited | None announced |
| Consumer access | Gemini app on AI Pro and Ultra, AI Mode in Search | Not yet; AI Ultra subscribers are in the first wave, no date |
A few cells need context. Google hasn’t published Argon’s input window, though its own long-context eval used prompts up to 1M tokens. Google says Argon’s output limit is 1M tokens; at least one third-party evaluator, Vals AI, lists a 262K max output for the configuration it tested. Argon’s evals ran with “the highest thinking settings,” so a high level likely exists, but Google hasn’t named the levels or the default. Flash’s levels are on Google’s thinking docs and in our 3.8 Flash thinking levels guide.
Flash’s free tier is rate-limited, and Google says free-tier data is “used to improve our products.” Google hasn’t announced any free way to use Argon; our Argon free access explainer covers what you can use instead.
The benchmark rows both launch posts share in text
Google’s launch tables for the two models share two rows in text. These are Google-reported numbers, and Argon’s methodology attributes both rows to Vals AI.
| Benchmark | Gemini 3.8 Flash (Sep 2 table) | Gemini 4 Argon (Sep 30 table) |
|---|---|---|
| Vals Finance Agent v2 | 61.4% | 65.4% |
| Harvey Legal Agent Benchmark | 10.0% | 19.6% |
Google published these tables a month apart, against different rival sets, so read the gap as directional rather than a controlled test. Even so, the legal row stands out: Argon’s 19.6% is nearly double Flash’s 10.0%. The finance gap is 4 points.
Everything else in Argon’s table has no 3.8 Flash counterpart in text. Google’s 3.8 Flash table reported HLE-Verified, which Argon’s doesn’t, and Flash’s DeepSWE score only appeared in model card images. On Arena’s Text leaderboard, a third-party ranking, Argon (High) sits at #1 on preliminary votes, while Gemini 3.8 Flash (High) is #11. The full 19-row Argon table, with who measured each row, is in our Argon benchmarks breakdown.
Cyber: Argon vs 3.8 Flash Cyber
The DeepMind cyber page compares Argon with Gemini 3.8 Flash Cyber, the Fairwind-gated cyber sibling of Flash:
| Cyber eval | Gemini 4 Argon | Gemini 3.8 Flash Cyber |
|---|---|---|
| Real-world vulnerability discovery | 85.8% | 71.0% |
| Wiz Penetration Test Benchmark | 70.9% | 58.2% |
Both models are gated. Gemini 3.8 Flash Cyber is GA behind an allowlist on Gemini Enterprise Agent Platform, and Argon is Fairwind-only. If you aren’t a Fairwind partner, neither number changes what you can call today. The general 3.8 Flash model is the one you can build on.
Cost of the same call on both
Take one call with 100,000 input tokens and 8,000 output tokens. Current Gemini models, 3.8 Flash included, bill thinking tokens as output, so the 8,000 includes them. Google hasn’t said how Argon bills them; the Argon rows assume it follows current Gemini models.
| Model and period | Input cost | Output cost | Total per call |
|---|---|---|---|
| 3.8 Flash, intro | 0.1M x $0.75 = $0.075 | 0.008M x $3.75 = $0.030 | $0.105 |
| 3.8 Flash, from 2027-01-01 | 0.1M x $1.50 = $0.150 | 0.008M x $7.50 = $0.060 | $0.210 |
| Argon, intro | 0.1M x $2 = $0.200 | 0.008M x $10 = $0.080 | $0.280 |
| Argon, standard | 0.1M x $4 = $0.400 | 0.008M x $20 = $0.160 | $0.560 |
Argon’s per-token rates are 8/3 (about 2.67 times) Flash’s, on both intro pricing and standard pricing. At 1,000 of these calls a day, that’s $105 a day on Flash against $280 on Argon at intro rates.
Three things break the tidy ratio:
- Token counts won’t match. The same prompt produces different output and thinking token counts on different models. Argon’s published evals ran at the highest thinking settings, and on 3.8 Flash Google warns that higher effort levels can use more tokens.
- Long outputs change the math. A 200,000-token answer doesn’t fit in Flash’s 65,536-token cap, so you’d split it across several calls. Under Argon’s stated limit it’s one response: 0.2M x $10 = $2.00 at intro rates, or $4.00 at standard. A full 1M-token answer costs $10 intro or $20 standard.
- Only one intro end date is known. Flash’s intro pricing ends 2026-12-31. Google hasn’t said when Argon’s ends.
Flash also has Batch at 50% off ($0.375 and $1.875 through December 31), per the Gemini API pricing page. Google hasn’t published Batch pricing for Argon. Our Argon pricing guide and 3.8 Flash pricing guide go deeper.
Should you wait for Argon or ship on Flash?
Ship on 3.8 Flash now when
- You’re cost-bound. Flash costs 3/8 of Argon per token, and Batch halves that.
- You run high volume. Classification, extraction, routing and chat at scale add up fast at $10 per 1M output tokens.
- You need it today. Flash has a stable model ID, a free tier for prototyping, and a published price schedule.
- Your responses fit in 65,536 tokens. Most API workloads do.
Plan for Argon when
- The work is long-horizon. Google says Argon is “built to sustain deep reasoning across complex, long-horizon workflows.”
- You need outputs beyond 64K. Full-module rewrites, long reports and large migrations are the case for a 1M output limit.
- You run legal or finance agents. Those are the two rows where Google published both models, and Argon leads both.
- You’re a Fairwind-eligible defender. Argon is the model the program now leads with.
You don’t have to pick once. Route the bulk of traffic to Flash and send the hard tail to Argon when it ships, with both behind the same config.
One saved request, two models
Here’s the pattern. Hold the model name in an environment variable, run your real requests against 3.8 Flash now, and swap the variable the day Argon’s model ID ships. Google hasn’t published that ID, so don’t guess one.
MODEL="${GEMINI_MODEL:-gemini-3.8-flash}"
curl -s -X POST \
"https://generativelanguage.googleapis.com/v1beta/models/${MODEL}:generateContent" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"contents":[{"parts":[{"text":"Summarize this contract clause in 3 bullets."}]}],
"generationConfig":{"thinkingConfig":{"thinkingLevel":"medium"},"maxOutputTokens":8192}}'
Every generateContent response ends with usageMetadata. In Apidog, store GEMINI_API_KEY and GEMINI_MODEL in an environment, save this request, and add assertions: status 200, the fields your app reads, and a ceiling on usageMetadata.thoughtsTokenCount sized to your cost cap. Keep maxOutputTokens set so a long answer can’t run away with your budget. Add it to a test scenario, run it on Flash and keep the report as a baseline.
Two caveats for launch day. Google says “all new models” will launch on the Interactions API and hasn’t said whether Argon supports generateContent, so save an Interactions version (POST https://generativelanguage.googleapis.com/v1beta/interactions with generation_config.thinking_level) next to this one. And Argon’s thinking level names aren’t documented, so medium may not carry over. When Argon lands, change one variable, rerun, and compare outputs and usageMetadata side by side.
FAQ
Is Gemini 4 Argon better than Gemini 3.8 Flash? On the two rows both launch tables share in text, yes: 65.4% vs 61.4% on Vals Finance Agent v2 and 19.6% vs 10.0% on Harvey’s Legal Agent Benchmark. It also costs about 2.67 times as much per token, and you can’t call it yet.
Should I wait for Gemini 4 Argon? Not if you need to ship. Google has given no release date, only “as soon as possible,” starting with paid API customers and AI Ultra subscribers.
Gemini 3.8 Flash or Argon for a new project? Start on 3.8 Flash with the model in a variable. Move the requests that need long outputs or legal and finance depth to Argon once you’ve tested it on your own prompts.
Is there a free way to use Argon? No. Google hasn’t announced a free tier; see our Argon free access explainer for the free Gemini models you can use today.
Does Argon replace Gemini 3.8 Flash? Google hasn’t said. Google calls Argon its new frontier model, and Reuters reports it’s larger than the previous Pro line, so it’s a different tier from Flash.
Next step
Set up the request today, run it on 3.8 Flash, and save the report. When Argon ships, you’ll know within minutes whether it earns its price on your traffic. Download Apidog to build the comparison.



