There is no public Gemini 4 Argon API yet, and Google hasn’t published a model ID. Google announced Argon on September 30, 2026, and it’s available today only to Fairwind Program defenders: a set of Fairwind partners use it as a managed model on Gemini Enterprise. Artificial Analysis lists Google AI Studio as Argon’s one API provider, but with no speed or latency data, which fits allowlisted pre-release access rather than an endpoint you can sign up for. When the wider rollout starts, Google says it begins “with paid API customers and Google AI Ultra subscribers,” so a paid Gemini API key puts you first in line.
That gap between announcement and access is time you can use. This guide separates what Google has confirmed about the Argon API from what it hasn’t, then walks through code you can run today on Gemini 3.8 Flash so the switch to Argon becomes a one-variable change. You’ll build the request, set up a go-live alert, mock the response in Apidog, and add a cost ceiling at Argon’s prices. For the model itself, start with what Gemini 4 Argon is; for rollout stages, see the Gemini 4 Argon release date guide.
What’s confirmed about the Gemini 4 Argon API, and what isn’t
Google’s launch post gives prices and an output limit. Almost everything else an integration needs is still unpublished.
| Item | Status | Detail |
|---|---|---|
| Input price | Confirmed | $2 per 1M tokens (intro), $4 after the intro period |
| Output price | Confirmed | $10 per 1M tokens (intro), $20 after |
| Cached input | Confirmed | 95% off input: $0.10 intro, $0.20 standard |
| Output limit | Confirmed | 1M tokens, up from 64K |
| API surface | Signaled | Google’s docs say all new models launch on the Interactions API |
| Model ID | Not published | Absent from the models page, pricing page, and changelog |
| Input context window | Not published | Google’s long-context eval used prompts up to 1M tokens, but that isn’t a spec |
| Thinking levels | Not published | Evals ran at the “highest thinking settings”; names and default unknown |
| Rate limits | Not published | No tiers announced |
| Batch support | Not published | No batch or Flex pricing |
| Long-prompt tier | Not published | 3.1 Pro charges more above 200K; Argon’s rule is unstated |
| Intro period length | Not published | No end date for $2/$10 |
Two rows deserve a closer look. The API surface row comes from the Interactions API docs, which say new models “will launch on the Interactions API.” Argon is a new model, so expect it there first; Google hasn’t said whether generateContent will serve it. The output row is Google’s stated limit for the model, but Vals AI lists a 262K max output for the configuration it tested, so don’t assume every endpoint serves the full million on day one. Our 1M output tokens guide covers what that limit does to streaming, timeouts, and storage.

On price, one paragraph is enough here. A request with a 20,000-token prompt and 5,000 output tokens costs 20,000 x $2/1M + 5,000 x $10/1M = $0.04 + $0.05 = $0.09 at intro rates, and $0.18 at the standard $4/$20. Argon’s intro output price ($10) sits below Gemini 3.1 Pro Preview’s $12, and its standard price matches Claude Opus 5.5 exactly. The Gemini 4 Argon pricing breakdown runs the full scenarios, including cached input and a maxed-out 1M-token response.
Don’t hard-code a model ID you found online
Search for the Argon model ID and you’ll find strings on benchmark sites, aggregators, and open-source pull requests. They’re placeholders. Vals AI uses its own slug for its model page, one GitHub pull request adds an ID “provisionally,” and the OpenRouter-style path leads to a not-found page. Google hasn’t confirmed any of them, and several outlets explicitly warn against hard-coding a guess.
A guessed ID fails in the least helpful way: a 404 in production on deploy day, or a silent fallback if your wrapper catches errors too broadly. Keep the model name in configuration instead. In every example below, the model comes from a GEMINI_MODEL environment variable that defaults to gemini-3.8-flash, a stable model you can call today. When Google publishes Argon’s ID, you change one value.
Get your code ready on Gemini 3.8 Flash
Gemini 3.8 Flash (gemini-3.8-flash) runs on the same endpoints, headers, and response shapes Argon is expected to use, at $0.75/$3.75 per 1M tokens through December 31, 2026. Your key lives in GEMINI_API_KEY; never paste it into code. Full setup is in our Gemini 3.8 Flash API guide.
Step 1: Send an Interactions API request
The Interactions API is Google’s primary API, Generally Available since June 2026. Keep both the model and the thinking level in variables, because Argon’s level names aren’t published.
MODEL="${GEMINI_MODEL:-gemini-3.8-flash}"
THINKING="${GEMINI_THINKING:-medium}"
curl -s -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"model\": \"${MODEL}\",
\"input\": \"List the retry rules a REST client should follow for HTTP 429.\",
\"generation_config\": {\"thinking_level\": \"${THINKING}\"}
}"
The Python SDK needs google-genai 2.3.0 or later for the Interactions API:
# pip install "google-genai>=2.3.0"
import os
from google import genai
MODEL = os.environ.get("GEMINI_MODEL", "gemini-3.8-flash")
THINKING = os.environ.get("GEMINI_THINKING", "medium")
client = genai.Client() # reads GEMINI_API_KEY from the environment
interaction = client.interactions.create(
model=MODEL,
input="List the retry rules a REST client should follow for HTTP 429.",
generation_config={"thinking_level": THINKING},
)
print(interaction.output_text)
On 3.8 Flash the valid levels are low, medium (the default), and high; minimal returns HTTP 400. Our thinking levels guide explains the trade-offs. For multi-turn work, pass the previous response’s ID as previous_interaction_id, and send store: false to opt out of server-side storage.
Step 2: Keep the generateContent path working
Most existing code calls the legacy generateContent endpoint, which Google says remains fully supported. Here’s the same request:
curl -s "https://generativelanguage.googleapis.com/v1beta/models/${MODEL}:generateContent" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"contents\": [{\"parts\": [{\"text\": \"List the retry rules a REST client should follow for HTTP 429.\"}]}],
\"generationConfig\": {\"thinkingConfig\": {\"thinkingLevel\": \"${THINKING}\"}}
}"
Note where thinking lives: generation_config.thinking_level on Interactions, generationConfig.thinkingConfig.thinkingLevel here. If your client wraps both, route new models through Interactions by default. Tool definitions need the same care; see Gemini 3.8 Flash function calling.
Step 3: Schedule a go-live check with models.list
Don’t refresh the changelog. Ask the API. The models endpoint returns the available models with their token limits and supported methods:
import os, sys, requests
resp = requests.get(
"https://generativelanguage.googleapis.com/v1beta/models",
params={"pageSize": 1000},
headers={"x-goog-api-key": os.environ["GEMINI_API_KEY"]},
timeout=30,
)
resp.raise_for_status()
hits = [m for m in resp.json().get("models", []) if "argon" in m["name"].lower()]
for m in hits:
print(m["name"], "in:", m.get("inputTokenLimit"),
"out:", m.get("outputTokenLimit"), m.get("supportedGenerationMethods"))
sys.exit(1 if hits else 0) # a failed run is the alert
Run it hourly from cron or a CI schedule, using the key your production app uses. When we ran this call on October 1, 2026, with a paid project key, it returned 61 models, no nextPageToken, and no name containing “argon”. The day it matches, the entry tells you three things at once: the real model name, whether outputTokenLimit is the full 1M, and which methods are listed.
Step 4: Mock the response so the client is built before access
You don’t need Argon to build the code that consumes Argon. Save the Step 2 request in Apidog with GEMINI_API_KEY and GEMINI_MODEL as environment variables, send it once against 3.8 Flash, and extract the real response into the endpoint as a response example. Apidog’s default Smart Mock generates data from the schema, so set the project’s default mock method to “Response example first” (Project Settings, Feature Settings, Mock Settings). The mock URL then returns your saved response, so your front end, queue workers, and parsers can run against it without spending tokens.
Be clear about what this mock is: the current Gemini response schema, not an Argon-specific one, because Google hasn’t published one. It’s still the right bet, since Argon is expected on the same API. To rehearse a heavy Argon reply, edit the example’s usageMetadata to large counts and confirm your billing and alerting code reacts.
Step 5: Assert on usageMetadata and a cost ceiling
Every generateContent response includes usageMetadata. Add a post-processor script in Apidog that prices each response at Argon’s standard rates, so a test fails before a bill does:
// Apidog post-processor: price this response at Gemini 4 Argon's standard rates
const u = pm.response.json().usageMetadata;
const IN = 4 / 1e6; // USD per input token after the intro period
const OUT = 20 / 1e6; // USD per output token; thinking bills as output on current Gemini models
const out = (u.candidatesTokenCount || 0) + (u.thoughtsTokenCount || 0);
const cost = u.promptTokenCount * IN + out * OUT;
pm.test("usageMetadata is present", () => pm.expect(u).to.be.an("object"));
pm.test("cost under $0.10 at Argon rates", () => pm.expect(cost).to.be.below(0.10));
Pick the ceiling per request; $0.10 suits a short prompt like this one. Google hasn’t stated how Argon bills thinking tokens, so the script assumes the current Gemini rule. Keep the same saved request and run it against 3.8 Flash now and Argon later: the difference in token counts and cost is your regression comparison.
FAQ
Is the Gemini 4 Argon API available? Not publicly. Argon is rolling out only to a set of Fairwind Program partners, who use it through Gemini Enterprise. Paid API customers and Google AI Ultra subscribers come next, with no date.
What is the Gemini 4 Argon model ID? Google hasn’t published one. Strings on third-party sites are placeholders, so keep the model in an environment variable and watch models.list.
How much will the Gemini 4 Argon API cost? $2 input and $10 output per 1M tokens during an intro period of unstated length, then $4 and $20. Cached input is 95% off. See Gemini 4 Argon pricing.
Will Argon work with generateContent? Google hasn’t said. Its docs say all new models launch on the Interactions API, so build on Interactions and treat generateContent as a fallback.
Does Google AI Ultra give me API access? No. AI Ultra is a consumer subscription, not an API key. Google says API access starts with paid API customers.
Your next step
Set GEMINI_MODEL in every environment, schedule the models.list check, and save the Interactions and generateContent requests with the cost assertion attached. Download Apidog to keep those requests, mocks, and tests in one workspace, then change one variable when Google publishes the ID.



