Gemini 4 Argon costs $2 per 1M input tokens and $10 per 1M output tokens during an introductory period, then $4 and $20. Cached input is 95% off the input price, which works out to $0.10 per 1M tokens at the intro rate and $0.20 after it. Google hasn’t said how long the intro period lasts, and you can’t buy Argon yet: it’s announced but not generally available, and today it’s rolling out only to Fairwind Program cyber defenders.
This post turns that price sheet into budget numbers: a comparison with seven current models, four worked scenarios, third-party cost per task, and the cost controls to set up now. For the model itself, start with what Gemini 4 Argon is; for access, see is Gemini 4 Argon free. You can build cost assertions in Apidog against a model you can call today and swap the model name later.

The price sheet
Google published the prices in the Gemini 4 Argon launch post, with the post-intro rate in footnote 1: “After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.”
| Gemini 4 Argon, per 1M tokens | Intro | Standard (after intro) |
|---|---|---|
| Input | $2.00 | $4.00 |
| Cached input (95% off input) | $0.10 | $0.20 |
| Output | $10.00 | $20.00 |
| Max output per response (Google) | 1M tokens | 1M tokens |
| Cost of one full 1M-token output | $10.00 | $20.00 |
The cached rates are arithmetic from Google’s 95% rule ($2 x 0.05 and $4 x 0.05). Google hasn’t published an input context window or a model ID. It says the output limit is 1M tokens, “up from the previous 64K tokens”; at least one third-party evaluator, Vals AI, lists a 262K max output for the configuration it tested.
Argon vs current frontier and Gemini models
All prices are per 1M tokens.
| Model | Input | Cached input | Output | Max output |
|---|---|---|---|---|
| Gemini 4 Argon (intro) | $2 | $0.10 | $10 | 1M |
| Gemini 4 Argon (standard) | $4 | $0.20 | $20 | 1M |
| GPT-6 Astra | $10 | $1 | $50 | 128,000 |
| Claude Opus 5.5 | $4 | $0.20 | $20 | 128K (300K on Batch, beta) |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 | 128K (300K on Batch, beta) |
| GPT-6.1 Sol | $2 | $0.10 | $10 | 128,000 |
| Claude Fable 5.1 | $10 | $0.25 | $50 | 128K |
| gemini-3.1-pro-preview (prompt <=200K / >200K) | $2 / $4 | $0.20 / $0.40 | $12 / $18 | 65,536 |
| gemini-3.8-flash (intro / from 2027-01-01) | $0.75 / $1.50 | $0.075 / $0.15 | $3.75 / $7.50 | 65,536 |
Competitor prices come from the vendors’ own pages: OpenAI’s GPT-6 Astra page, Anthropic’s pricing page, and Google’s Gemini API pricing page.
- Standard Argon = Claude Opus 5.5. Same $4/$20, same $0.20 cached input.
- Intro Argon = Claude Sonnet 5.5 and GPT-6.1 Sol. Same $2/$10; the $0.10 intro cached rate matches GPT-6.1 Sol.
- GPT-6 Astra and Claude Fable 5.1 cost 5x intro Argon and 2.5x standard Argon. See Claude Fable 5.1 pricing.
- Argon’s intro output ($10) is below gemini-3.1-pro-preview’s ($12), Google’s previous top Pro model.
- Long prompts change the math. OpenAI bills prompts over 272K input tokens at 2x input and cache and 1.5x output for the full request. Gemini 3.1 Pro charges more above 200K; Anthropic doesn’t. Google hasn’t said whether Argon has a long-prompt tier.
For capability alongside price, see Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5.
Thinking tokens bill as output
On current Gemini models, thinking tokens bill as output: Google’s thinking docs say “response pricing is the sum of output tokens and thinking tokens.” Google hasn’t documented Argon specifically, but plan on the same rule. Google ran Argon’s evals “with the highest thinking settings,” so read “$10 per 1M output tokens” as answer plus reasoning.
Four worked scenarios
Every figure below is tokens divided by 1M, times the per-1M rate. Scenarios 1 to 3 assume no caching.
1. A typical API call: 20K in, 5K out
- Intro: 20,000 / 1M x $2 = $0.04 input, plus 5,000 / 1M x $10 = $0.05 output. Total $0.09.
- Standard: $0.08 + $0.10 = $0.18.
At 1,000 calls a day, that’s $90 intro or $180 standard. The same call costs $0.18 on Claude Opus 5.5, $0.45 on GPT-6 Astra ($0.20 + $0.25), $0.10 on gemini-3.1-pro-preview ($0.04 + $0.06), and about $0.034 on gemini-3.8-flash at its intro rate ($0.015 + $0.019).
2. A long agent turn: 200K in, 50K out
- Intro: 200,000 / 1M x $2 = $0.40, plus 50,000 / 1M x $10 = $0.50. Total $0.90.
- Standard: $0.80 + $1.00 = $1.80.
GPT-6 Astra would charge $4.50 ($2.00 + $2.50). gemini-3.1-pro-preview charges $1.00 ($0.40 + $0.60), but 200K is its tier line: longer prompts bill at $4/$18. If Argon gets a similar tier, longer turns cost more.
3. One maxed-out response: 1M tokens of output
Output alone is 1,000,000 / 1M x $10 = $10.00 at intro and $20.00 at standard. Add a 100K-token prompt ($0.20 intro, $0.40 standard) and the call costs $10.20 or $20.40.
No other model in the table produces this in one synchronous response: competitors cap at 128K, the previous Gemini Pro at 65,536. At Vals’ listed 262K max output, a full response costs about $2.62 intro or $5.24 standard. The engineering side (timeouts, streaming, gateways) is in Gemini 4 Argon’s 1M output tokens.
4. A cached 500K-token prompt reused 10 times
Say you send the same 500K-token codebase or contract set ten times, with 5K tokens of output each. Assume the first call pays full input price to build the cache and the next nine read from it.
| 10 calls, 500K prompt, 5K output each | Intro | Standard |
|---|---|---|
| Input, no caching (10 x 500K) | $10.00 | $20.00 |
| Input, cached (1 full + 9 cached) | $1.00 + 9 x $0.05 = $1.45 | $2.00 + 9 x $0.10 = $2.90 |
| Output (10 x 5K) | $0.50 | $1.00 |
| Total with caching | $1.95 | $3.90 |
| Total without caching | $10.50 | $21.00 |
Caching cuts the input bill by 85.5%. Two caveats: Google hasn’t said whether Argon charges for cache storage (gemini-3.1-pro-preview charges $4.50 per 1M tokens per hour, which would add $2.25 to hold 500K tokens for an hour), and a 500K prompt is past 3.1 Pro’s 200K line. Argon’s standard cached rate matches Opus 5.5’s $0.20, so the break-even logic in Claude Opus 5.5 prompt caching cost math carries over.
Per-token price is not per-task cost
Artificial Analysis lists $1.99 per task to run its Intelligence Index on Gemini 4 Argon (High) at intro prices; The Decoder puts the same run at $3.98 at standard prices. AA scores Argon 53, the same as GPT-6 Astra (max), which it lists at $3.26 per task.
The catch is token volume. AA counts roughly 62K output tokens per task for Argon (high) against roughly 27K for Astra at its max setting, about 2.3x. Output per task at intro: 62,000 / 1M x $10 = $0.62, vs 27,000 / 1M x $50 = $1.35 for Astra. At standard, The Decoder’s full-task figure for Argon ($3.98) lands above AA’s $3.26 for Astra (max), even though Argon’s standard rates are 60% below Astra’s.
Vals AI shows the same pattern. It lists $15.68 per test for Argon on the Vals Index, against $32.14 for Claude Opus 5.5, $21.34 for Claude Sonnet 5.5, and $3.24 for GPT-6.1 Sol, and it lists Argon at the $4/$20 standard rate. Sonnet 5.5 and GPT-6.1 Sol share Argon’s intro price on paper, yet their per-test costs land more than 6x apart.
Budget from your own token counts, not the rate card.
Cost controls to set up before Argon ships
- Cap output. On generateContent,
generationConfig.maxOutputTokenssets a ceiling. Google says new models launch on the Interactions API, so set the equivalent cap there once the docs list it for Argon. A 1M ceiling is a $20 output ceiling per call at standard rates. - Cache the stable content. System instructions, schemas, and reference documents that repeat across calls are where the 95% discount pays off.
- Choose the thinking level on purpose. Google hasn’t published Argon’s levels. Today gemini-3.1-pro-preview defaults to
highand gemini-3.8-flash tomedium. When Argon’s are documented, test whether a lower one holds quality on your task. - Assert on usage. Every generateContent response ends with
usageMetadata. In Apidog, save the request with the model in aGEMINI_MODELenvironment variable, add a post-response assertion that computes cost from the token counts inusageMetadata, and fail the test when a call crosses your ceiling. Run it against gemini-3.8-flash today; when Argon’s ID ships, change one variable and rerun the suite.
The arithmetic for that assertion:
cost = uncached_input / 1,000,000 x input_rate
+ cached_input / 1,000,000 x cached_rate
+ (output + thinking) / 1,000,000 x output_rate
What Google hasn’t priced yet
The intro period’s length and the date $4/$20 starts; whether prompts above 200K cost more; Batch or Flex pricing (3.1 Pro has both, at $1/$2 input and $6/$9 output); cache storage fees; rate limits; and whether there’s a free tier at all.
FAQ
How much does Gemini 4 Argon cost per million tokens? $2 input and $10 output during the intro period, then $4 and $20. Cached input is 95% off: $0.10, then $0.20.
How long does the intro price last? Google hasn’t said. The launch post gives no end date or length.
Are thinking tokens billed as output? On current Gemini models, yes. Google hasn’t documented Argon’s rules specifically.
How much does a 1M-token output response cost? $10 at intro pricing and $20 at standard, before input.
Is Argon cheaper than Claude Opus 5.5 or GPT-6 Astra? Per token, standard Argon equals Opus 5.5 and sits 60% below Astra. Per task, it depends on token volume: Artificial Analysis measured Argon using about 2.3x Astra’s output tokens. For a cheaper Gemini today, see Gemini 3.8 Flash pricing.
Can I pay for Argon today? No. It’s rolling out only to Fairwind Program partners. Paid API customers and Google AI Ultra subscribers come next, with no date.
Budget it now, measure it later
Multiply your current traffic’s token counts by $2/$10 and $4/$20, and you have Argon’s intro and standard range. Then build the measurement: download Apidog, save your request against gemini-3.8-flash with GEMINI_MODEL as a variable, and add a cost assertion on usageMetadata. When Google publishes the model ID, you swap one value and see your real per-call cost on day one.



