Gemini 4 Argon's 1M Output Tokens: What a Million-Token Response Does to Your API Stack

Gemini 4 Argon's 1M output tokens is an output limit, not a context window. What a maxed response costs, plus streaming, timeouts, and caps.

Ashley Goolam

Ashley Goolam

2 October 2026

Gemini 4 Argon's 1M Output Tokens: What a Million-Token Response Does to Your API Stack

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Gemini 4 Argon’s headline number, 1M tokens, is its output limit, not its context window. Google says a single Argon response can run to 1 million tokens, roughly 16x the previous 64K cap, and it hasn’t published Argon’s input window at all. There’s also nothing to call yet: Argon is available today only to Fairwind Program defenders, with paid API customers next once Google opens access (see the release date and access guide).

A response that long breaks three assumptions most API stacks rely on: a call finishes in seconds, the body fits in memory, and one request has a small, predictable cost. This guide covers the cost of a maxed-out response, streaming, timeouts, output caps, storage, and how to test all of it in Apidog before access arrives. For the model overview, see what Gemini 4 Argon is; for request shapes, see the Gemini 4 Argon API guide.

1M output tokens is not a 1M context window

Several pages ranking for Argon describe the 1M figure as a context window, and one headline calls it a “16-fold larger context window.” That gets it backwards. Google’s launch post says it expanded “the model’s output token limit to an industry-leading 1M tokens.” That 64K matches the 65,536-token output cap on Gemini 3.1 Pro Preview, Google’s previous top Pro model.

Input is a separate number Google hasn’t given. Its long-context eval, described in the evals methodology, used prompts between 256K and 1M tokens. That’s a benchmark subset, not a spec. There’s a caveat on the output side too: Vals AI lists a 262K max output for the Argon configuration it tested. Google says the model’s limit is 1M; at least one third-party evaluator saw a lower cap on the endpoint it used.

Model Max output per response Input or context window
Gemini 4 Argon 1M (Google’s stated limit) Not published
Gemini 3.1 Pro Preview 65,536 1,048,576
Gemini 3.8 Flash 65,536 1,048,576
GPT-6 Astra 128,000 1,050,000 (922K max input)
Claude Opus 5.5 128K (300K on Batch with a beta header) 1M

Every competitor caps synchronous output at 128K, so Argon’s stated limit is about 8x theirs. Google’s reason is depth of reasoning: with headroom, the model can “generate hundreds of thousands of tokens in a single trajectory” and solve hard problems in one pass. For how other vendors handle long runs, see Claude Opus 5.5’s 18-hour tasks and the GPT-6 Astra API guide.

What one maxed-out response costs

Price the ceiling first. Output bills at $10 per 1M tokens during Argon’s intro period and $20 after, so one full-length response costs:

Input comes on top. A 200,000-token prompt adds 200,000 x $2/1M = $0.40 at intro rates or $0.80 at standard, so a single maxed call runs $10.40 or $20.80. A nightly job that fires 100 of those costs $1,040 at intro rates.

Thinking makes this harder to see. On current Gemini models, thinking tokens bill as output; Google hasn’t said whether Argon follows that rule or whether thinking counts toward the 1M ceiling. Either way, a short visible answer can still carry a large output bill. The Gemini 4 Argon pricing guide runs more scenarios, including cached input at 95% off.

Why streaming is mandatory

A non-streaming call returns nothing until the whole response is done. At hundreds of thousands of tokens, that’s a long silent connection, and idle timeouts across your stack can close it before the first byte arrives.

Stream instead. On generateContent, swap the method for :streamGenerateContent?alt=sse and Google sends server-sent events, one partial candidates chunk per event. Read each event as it lands and write it out; don’t collect the body first. This runs on Gemini 3.8 Flash today, with the model in a variable because Google hasn’t published Argon’s model ID (setup is in our Gemini 3.8 Flash API guide):

import json, os, requests

MODEL = os.environ.get("GEMINI_MODEL", "gemini-3.8-flash")
URL = ("https://generativelanguage.googleapis.com/v1beta/models/"
       f"{MODEL}:streamGenerateContent?alt=sse")
body = {
    "contents": [{"parts": [{"text": "Write a test plan for every endpoint in a payments API."}]}],
    "generationConfig": {"maxOutputTokens": 60000},
}
usage = None
with requests.post(URL, json=body, stream=True, timeout=(10, 120),
                   headers={"x-goog-api-key": os.environ["GEMINI_API_KEY"]}) as r, \
        open("response.txt", "a", encoding="utf-8") as out:
    r.raise_for_status()
    for line in r.iter_lines(decode_unicode=True):
        if not line or not line.startswith("data:"):
            continue
        event = json.loads(line[5:])
        for cand in event.get("candidates", []):
            for part in cand.get("content", {}).get("parts", []):
                out.write(part.get("text", ""))
        out.flush()
        usage = event.get("usageMetadata", usage)
print(usage)

timeout=(10, 120) sets a 10-second connect timeout and a 120-second read timeout. In requests, the read timeout is the longest gap between bytes, not the total duration, so a stream that keeps sending can run as long as it needs. Each chunk goes to disk as it arrives. On 3.8 Flash every event carries a running usageMetadata, so the last one gives you the final token counts to log and bill against.

Timeouts at every hop

Your client is one hop. A long stream also crosses a reverse proxy, an API gateway, a load balancer, and maybe a serverless runtime, and any of them can end the response early:

Hop What to check Symptom when it’s wrong
HTTP client Read or idle timeout, plus any total-request timeout Exceptions mid-stream on long answers only
Reverse proxy Read timeout and response buffering for text/event-stream Events arrive in bursts, or the stream cuts off
API gateway Maximum request duration Requests fail at the same elapsed time every run
Load balancer Idle timeout Drops during long pauses before the first event
Serverless function Maximum execution time The function ends while the model is still writing

Watch for a fixed cutoff. If long responses always fail at the same elapsed time, some hop has a hard duration limit that streaming can’t fix, and that work needs to move off the request path.

Run long jobs in the background

For the longest jobs, take the work off a live connection. The Interactions API supports background execution for long-running tasks using background=true. Background runs depend on stored interactions: the docs say store=false is incompatible with background execution, so leave storage on for these requests. For retrieving a finished background interaction, follow Google’s docs; don’t guess at polling endpoints. Since Google says new models launch on the Interactions API, plan for Argon’s long jobs to run there.

Cap output on purpose

The 1M limit is a ceiling, not a target. On generateContent, generationConfig.maxOutputTokens caps each response; the streaming example sets 60,000. On 3.8 Flash, thinking counts against that cap: our test run capped at 2,000 returned 1,340 thought tokens and 656 visible ones. For the Interactions API, confirm the output-cap field in Google’s docs before you rely on one. Then pick the cap from the cost you’ll accept per call:

Output cap Worst-case output cost, standard ($20/1M) Intro ($10/1M)
64,000 64,000 x $20/1M = $1.28 $0.64
128,000 $2.56 $1.28
500,000 $10.00 $5.00
1,000,000 $20.00 $10.00

A response that hits the cap stops early, so treat it as incomplete. Check the final event’s finishReason: MAX_TOKENS means the cap cut it off. Then either continue in a follow-up turn or raise the cap for that one job.

Store and parse huge outputs without buffering

A million tokens is megabytes of text per response. A few rules keep that from taking down a worker:

Test it in Apidog before access opens

You can rehearse all of this against a stand-in. Download Apidog and work through three checks.

Watch the stream. Send the streaming request against 3.8 Flash with GEMINI_API_KEY and GEMINI_MODEL as environment variables. Apidog parses text/event-stream responses and shows each event in its Timeline view as it arrives, so you can see chunk sizes, gaps, and the final usageMetadata.

Stream a far longer fake response. Real 3.8 Flash output tops out at 65,536 tokens, so run a local mock that streams far more in the same event shape:

# long_stream_mock.py: Gemini-shaped SSE for parser and timeout tests (fake data)
import json, time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

EVENTS, DELAY, CHUNK = 20000, 0.005, "lorem ipsum " * 40

class Handler(BaseHTTPRequestHandler):
    def do_POST(self):
        self.rfile.read(int(self.headers.get("Content-Length", 0)))
        self.send_response(200)
        self.send_header("Content-Type", "text/event-stream")
        self.end_headers()
        for i in range(EVENTS):
            event = {"candidates": [{"content": {"parts": [{"text": CHUNK}]}}]}
            if i == EVENTS - 1:  # fake counts sized like a near-max reply
                event["usageMetadata"] = {"promptTokenCount": 1200,
                    "candidatesTokenCount": 950000, "thoughtsTokenCount": 40000,
                    "totalTokenCount": 991200}
            self.wfile.write(f"data: {json.dumps(event)}\n\n".encode())
            self.wfile.flush()
            time.sleep(DELAY)

ThreadingHTTPServer(("127.0.0.1", 8787), Handler).serve_forever()

Point a “mock” environment’s base URL at http://127.0.0.1:8787 and send the same streaming request through it. The stream runs for about two minutes (128 seconds in our test) and carries 9.6 million characters of text, enough to expose a buffering parser, a proxy that holds events, or a timeout set too short.

Assert on token counts. On a non-streaming generateContent request, assert that candidatesTokenCount plus thoughtsTokenCount in usageMetadata stays at or under your cap and that the computed cost stays under your ceiling at Argon prices. The Argon API guide has a ready-made cost script.

FAQ

Is 1M Gemini 4 Argon’s context window? No. 1M is the output limit per response, up from 64K. Google hasn’t published Argon’s input window.

How much does a 1M-token Argon response cost? $10 of output at intro rates and $20 at standard, plus input. See Gemini 4 Argon pricing for more scenarios.

Can I generate a 1M-token response today? Not unless your organization is in the Fairwind cohort with Argon access. Gemini 3.8 Flash and 3.1 Pro Preview cap output at 65,536 tokens, and Vals AI lists 262K max output for the Argon config it tested.

Do I have to stream long Argon responses? Google hasn’t published streaming guidance for Argon, but a non-streaming call that runs for minutes is exposed to every idle timeout in your stack. Stream it, or use background execution on the Interactions API.

How does Argon’s output limit compare to GPT-6 Astra and Claude Opus 5.5? Both cap synchronous output at 128K; Anthropic allows 300K on Batch with a beta header. Argon’s stated 1M is about 8x that.

Your next step

Add streaming and an output cap to your Gemini client now, on 3.8 Flash, and run it against the long mock stream until nothing in your stack cuts it off. When Argon’s ID ships, change GEMINI_MODEL and rerun the same tests in Apidog.

Explore more

Gemini 4 Argon API: What's Confirmed, What It Will Cost, and How to Get Your Code Ready

Gemini 4 Argon API: What's Confirmed, What It Will Cost, and How to Get Your Code Ready

Gemini 4 Argon API: no public access or model ID yet. What Google confirmed on price and output, and code to get ready on Gemini 3.8 Flash today.

2 October 2026

Codex after DevDay 2026: cloud tasks, a voice CLI, /agents, Code Review, and Security Cloud

Codex after DevDay 2026: cloud tasks, a voice CLI, /agents, Code Review, and Security Cloud

Codex DevDay 2026 updates: reusable cloud environments, a voice CLI with /agents, Code Review, Security Cloud, and GPT-6.1 Sol as the new default model.

2 October 2026

Sign in with ChatGPT for developers: the OAuth flow, plan usage, and what it means for your API bill

Sign in with ChatGPT for developers: the OAuth flow, plan usage, and what it means for your API bill

Sign in with ChatGPT for developers: what your app receives, how Plus and Pro plan usage and weekly caps work, and how to test the OAuth flow.

2 October 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

Gemini 4 Argon's 1M Output Tokens: What a Million-Token Response Does to Your API Stack