Gemini 4 Argon’s headline number, 1M tokens, is its output limit, not its context window. Google says a single Argon response can run to 1 million tokens, roughly 16x the previous 64K cap, and it hasn’t published Argon’s input window at all. There’s also nothing to call yet: Argon is available today only to Fairwind Program defenders, with paid API customers next once Google opens access (see the release date and access guide).
A response that long breaks three assumptions most API stacks rely on: a call finishes in seconds, the body fits in memory, and one request has a small, predictable cost. This guide covers the cost of a maxed-out response, streaming, timeouts, output caps, storage, and how to test all of it in Apidog before access arrives. For the model overview, see what Gemini 4 Argon is; for request shapes, see the Gemini 4 Argon API guide.
1M output tokens is not a 1M context window
Several pages ranking for Argon describe the 1M figure as a context window, and one headline calls it a “16-fold larger context window.” That gets it backwards. Google’s launch post says it expanded “the model’s output token limit to an industry-leading 1M tokens.” That 64K matches the 65,536-token output cap on Gemini 3.1 Pro Preview, Google’s previous top Pro model.
Input is a separate number Google hasn’t given. Its long-context eval, described in the evals methodology, used prompts between 256K and 1M tokens. That’s a benchmark subset, not a spec. There’s a caveat on the output side too: Vals AI lists a 262K max output for the Argon configuration it tested. Google says the model’s limit is 1M; at least one third-party evaluator saw a lower cap on the endpoint it used.
| Model | Max output per response | Input or context window |
|---|---|---|
| Gemini 4 Argon | 1M (Google’s stated limit) | Not published |
| Gemini 3.1 Pro Preview | 65,536 | 1,048,576 |
| Gemini 3.8 Flash | 65,536 | 1,048,576 |
| GPT-6 Astra | 128,000 | 1,050,000 (922K max input) |
| Claude Opus 5.5 | 128K (300K on Batch with a beta header) | 1M |
Every competitor caps synchronous output at 128K, so Argon’s stated limit is about 8x theirs. Google’s reason is depth of reasoning: with headroom, the model can “generate hundreds of thousands of tokens in a single trajectory” and solve hard problems in one pass. For how other vendors handle long runs, see Claude Opus 5.5’s 18-hour tasks and the GPT-6 Astra API guide.

What one maxed-out response costs
Price the ceiling first. Output bills at $10 per 1M tokens during Argon’s intro period and $20 after, so one full-length response costs:
- Intro: 1,000,000 x $10/1M = $10.00 of output
- Standard: 1,000,000 x $20/1M = $20.00 of output
Input comes on top. A 200,000-token prompt adds 200,000 x $2/1M = $0.40 at intro rates or $0.80 at standard, so a single maxed call runs $10.40 or $20.80. A nightly job that fires 100 of those costs $1,040 at intro rates.
Thinking makes this harder to see. On current Gemini models, thinking tokens bill as output; Google hasn’t said whether Argon follows that rule or whether thinking counts toward the 1M ceiling. Either way, a short visible answer can still carry a large output bill. The Gemini 4 Argon pricing guide runs more scenarios, including cached input at 95% off.
Why streaming is mandatory
A non-streaming call returns nothing until the whole response is done. At hundreds of thousands of tokens, that’s a long silent connection, and idle timeouts across your stack can close it before the first byte arrives.
Stream instead. On generateContent, swap the method for :streamGenerateContent?alt=sse and Google sends server-sent events, one partial candidates chunk per event. Read each event as it lands and write it out; don’t collect the body first. This runs on Gemini 3.8 Flash today, with the model in a variable because Google hasn’t published Argon’s model ID (setup is in our Gemini 3.8 Flash API guide):
import json, os, requests
MODEL = os.environ.get("GEMINI_MODEL", "gemini-3.8-flash")
URL = ("https://generativelanguage.googleapis.com/v1beta/models/"
f"{MODEL}:streamGenerateContent?alt=sse")
body = {
"contents": [{"parts": [{"text": "Write a test plan for every endpoint in a payments API."}]}],
"generationConfig": {"maxOutputTokens": 60000},
}
usage = None
with requests.post(URL, json=body, stream=True, timeout=(10, 120),
headers={"x-goog-api-key": os.environ["GEMINI_API_KEY"]}) as r, \
open("response.txt", "a", encoding="utf-8") as out:
r.raise_for_status()
for line in r.iter_lines(decode_unicode=True):
if not line or not line.startswith("data:"):
continue
event = json.loads(line[5:])
for cand in event.get("candidates", []):
for part in cand.get("content", {}).get("parts", []):
out.write(part.get("text", ""))
out.flush()
usage = event.get("usageMetadata", usage)
print(usage)
timeout=(10, 120) sets a 10-second connect timeout and a 120-second read timeout. In requests, the read timeout is the longest gap between bytes, not the total duration, so a stream that keeps sending can run as long as it needs. Each chunk goes to disk as it arrives. On 3.8 Flash every event carries a running usageMetadata, so the last one gives you the final token counts to log and bill against.
Timeouts at every hop
Your client is one hop. A long stream also crosses a reverse proxy, an API gateway, a load balancer, and maybe a serverless runtime, and any of them can end the response early:
| Hop | What to check | Symptom when it’s wrong |
|---|---|---|
| HTTP client | Read or idle timeout, plus any total-request timeout | Exceptions mid-stream on long answers only |
| Reverse proxy | Read timeout and response buffering for text/event-stream |
Events arrive in bursts, or the stream cuts off |
| API gateway | Maximum request duration | Requests fail at the same elapsed time every run |
| Load balancer | Idle timeout | Drops during long pauses before the first event |
| Serverless function | Maximum execution time | The function ends while the model is still writing |
Watch for a fixed cutoff. If long responses always fail at the same elapsed time, some hop has a hard duration limit that streaming can’t fix, and that work needs to move off the request path.
Run long jobs in the background
For the longest jobs, take the work off a live connection. The Interactions API supports background execution for long-running tasks using background=true. Background runs depend on stored interactions: the docs say store=false is incompatible with background execution, so leave storage on for these requests. For retrieving a finished background interaction, follow Google’s docs; don’t guess at polling endpoints. Since Google says new models launch on the Interactions API, plan for Argon’s long jobs to run there.
Cap output on purpose
The 1M limit is a ceiling, not a target. On generateContent, generationConfig.maxOutputTokens caps each response; the streaming example sets 60,000. On 3.8 Flash, thinking counts against that cap: our test run capped at 2,000 returned 1,340 thought tokens and 656 visible ones. For the Interactions API, confirm the output-cap field in Google’s docs before you rely on one. Then pick the cap from the cost you’ll accept per call:
| Output cap | Worst-case output cost, standard ($20/1M) | Intro ($10/1M) |
|---|---|---|
| 64,000 | 64,000 x $20/1M = $1.28 | $0.64 |
| 128,000 | $2.56 | $1.28 |
| 500,000 | $10.00 | $5.00 |
| 1,000,000 | $20.00 | $10.00 |
A response that hits the cap stops early, so treat it as incomplete. Check the final event’s finishReason: MAX_TOKENS means the cap cut it off. Then either continue in a follow-up turn or raise the cap for that one job.
Store and parse huge outputs without buffering
A million tokens is megabytes of text per response. A few rules keep that from taking down a worker:
- Write chunks as they arrive, to a file or a multipart object upload, instead of building one string in memory.
- Keep the partial file if the connection drops. A blind retry from scratch generates, and bills, the output again.
- Ask for JSON Lines when you need structure, so each line parses on its own instead of waiting for one giant document to close.
- Log
usageMetadataand a byte count, not full bodies. - Check column and message-size limits in your database and queue before an 800K-token answer hits them.
Test it in Apidog before access opens
You can rehearse all of this against a stand-in. Download Apidog and work through three checks.
Watch the stream. Send the streaming request against 3.8 Flash with GEMINI_API_KEY and GEMINI_MODEL as environment variables. Apidog parses text/event-stream responses and shows each event in its Timeline view as it arrives, so you can see chunk sizes, gaps, and the final usageMetadata.
Stream a far longer fake response. Real 3.8 Flash output tops out at 65,536 tokens, so run a local mock that streams far more in the same event shape:
# long_stream_mock.py: Gemini-shaped SSE for parser and timeout tests (fake data)
import json, time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
EVENTS, DELAY, CHUNK = 20000, 0.005, "lorem ipsum " * 40
class Handler(BaseHTTPRequestHandler):
def do_POST(self):
self.rfile.read(int(self.headers.get("Content-Length", 0)))
self.send_response(200)
self.send_header("Content-Type", "text/event-stream")
self.end_headers()
for i in range(EVENTS):
event = {"candidates": [{"content": {"parts": [{"text": CHUNK}]}}]}
if i == EVENTS - 1: # fake counts sized like a near-max reply
event["usageMetadata"] = {"promptTokenCount": 1200,
"candidatesTokenCount": 950000, "thoughtsTokenCount": 40000,
"totalTokenCount": 991200}
self.wfile.write(f"data: {json.dumps(event)}\n\n".encode())
self.wfile.flush()
time.sleep(DELAY)
ThreadingHTTPServer(("127.0.0.1", 8787), Handler).serve_forever()
Point a “mock” environment’s base URL at http://127.0.0.1:8787 and send the same streaming request through it. The stream runs for about two minutes (128 seconds in our test) and carries 9.6 million characters of text, enough to expose a buffering parser, a proxy that holds events, or a timeout set too short.
Assert on token counts. On a non-streaming generateContent request, assert that candidatesTokenCount plus thoughtsTokenCount in usageMetadata stays at or under your cap and that the computed cost stays under your ceiling at Argon prices. The Argon API guide has a ready-made cost script.
FAQ
Is 1M Gemini 4 Argon’s context window? No. 1M is the output limit per response, up from 64K. Google hasn’t published Argon’s input window.
How much does a 1M-token Argon response cost? $10 of output at intro rates and $20 at standard, plus input. See Gemini 4 Argon pricing for more scenarios.
Can I generate a 1M-token response today? Not unless your organization is in the Fairwind cohort with Argon access. Gemini 3.8 Flash and 3.1 Pro Preview cap output at 65,536 tokens, and Vals AI lists 262K max output for the Argon config it tested.
Do I have to stream long Argon responses? Google hasn’t published streaming guidance for Argon, but a non-streaming call that runs for minutes is exposed to every idle timeout in your stack. Stream it, or use background execution on the Interactions API.
How does Argon’s output limit compare to GPT-6 Astra and Claude Opus 5.5? Both cap synchronous output at 128K; Anthropic allows 300K on Batch with a beta header. Argon’s stated 1M is about 8x that.
Your next step
Add streaming and an output cap to your Gemini client now, on 3.8 Flash, and run it against the long mock stream until nothing in your stack cuts it off. When Argon’s ID ships, change GEMINI_MODEL and rerun the same tests in Apidog.



