How to use GPT-6.1 Sol APl ?

GPT-6.1 Sol API guide: your first gpt-6.1-sol request, effort levels, Batch/Flex/Fast pricing, and the four changes to migrate from gpt-6-sol.

INEZA Felin-Michel

INEZA Felin-Michel

30 September 2026

How to use GPT-6.1 Sol APl ?

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

To call the GPT-6.1 Sol API, send a POST request to https://api.openai.com/v1/responses with "model": "gpt-6.1-sol" and your key as a Bearer token. It lists at the same $2 input and $10 output per million tokens as GPT-6 Sol, and cached input drops from $0.20 to $0.10. Migrating from gpt-6-sol is mostly a string swap. The breaking change is effort: GPT-6.1 Sol doesn’t accept none or minimal, so those requests move to low, along with any code that relied on none.

OpenAI shipped GPT-6.1 Sol at DevDay on September 29, 2026. The DevDay 2026 roundup covers the other launches, and what is GPT-6.1 Sol covers the benchmarks in depth. This guide covers your first request, which effort level to start at, every migration change, the Batch, Flex and Fast tiers, and a side-by-side regression run of both model IDs in Apidog before you switch production traffic.

button

GPT-6 Sol vs GPT-6.1 Sol: what changes in the API

Most of the spec is identical. Here’s the full diff from the GPT-6.1 Sol model page, the GPT-6 Sol model page and OpenAI’s Using GPT-6 migration guidance:

gpt-6-sol gpt-6.1-sol What to do
Input / output per 1M (Standard) $2 / $10 $2 / $10 Nothing
Cached input per 1M $0.20 $0.10 Re-run your cache math
Cache writes per 1M $2.50 $2.50 Nothing
Context window / max input / max output 1,050,000 / 922,000 / 128,000 1,050,000 / 922,000 / 128,000 Nothing
Knowledge cutoff Apr 20, 2026 Apr 30, 2026 Re-check date-sensitive evals
reasoning.effort none, low, medium (default), high, xhigh, max low, medium (default), high, xhigh, max Move none to low and re-evaluate
Function calling in Chat Completions Only with reasoning_effort: "none" Not supported Move tool calls to Responses
Endpoints Chat Completions, Responses, Batch Same Nothing
Rate limits Tier 1: 500 RPM / 500K TPM; Tier 5: 15,000 RPM / 40M TPM Same Nothing

The GPT-6 Sol page now sends readers to GPT-6.1 Sol as “the newer Sol model.”

Send your first GPT-6.1 Sol request

Export your key as OPENAI_API_KEY, then call the Responses API:

curl https://api.openai.com/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d '{
    "model": "gpt-6.1-sol",
    "reasoning": {"effort": "medium"},
    "input": "List three ways a webhook retry policy can create duplicate orders. One line each."
  }'

The Python SDK reads the same environment variable:

from openai import OpenAI

client = OpenAI()
response = client.responses.create(
    model="gpt-6.1-sol",
    reasoning={"effort": "medium"},
    input="List three ways a webhook retry policy can create duplicate orders. One line each.",
)
print(response.output_text)
print(response.usage)

Four parts of the response matter:

Use the Responses API for anything with tools; GPT-6.1 Sol supports Chat Completions only for requests without tools. The Responses API guide covers the request shape in more depth.

Choose a reasoning effort level

Effort is your main cost and quality dial, and medium is the default when you omit it. OpenAI’s model selection guide pairs medium with “complex technical work and coordinated deliverables you expect to revise,” and xhigh with polished deliverables and decisions built from conflicting evidence. OpenAI’s launch post adds results by setting. These benchmarks are OpenAI-reported, and the table quotes the deltas OpenAI states in its text:

Effort Start here for What OpenAI reports for GPT-6.1 Sol
low Chat, extraction, classification, anything you ran at none On user-flagged conversations, responses with a factual error fall from 11.4% (GPT-6 Sol) to 7.7%
medium (default) Agentic automations and tool-calling workflows AutomationBench 1.0.6: +2.2 pp over Claude Opus 5.5 at roughly a third of the cost; +4.8 pp over GPT-6 Sol at the same setting
high Hard debugging and deep planning No setting-specific claim
xhigh Polished deliverables and long asynchronous runs No setting-specific claim
max Computer use and hard science tasks OSWorld 2.0: +7 pp over GPT-6 Sol at max for less than half the cost. Terminal-Bench Science 0.1: $5.47 per task, vs $23.21 for Opus 5.5 and $23.80 for GPT-6 Astra

Two caveats. The factuality set is conversations previously flagged for errors, not typical traffic. And on Terminal-Bench Science, GPT-6 Astra still scores highest (68.1%), so OpenAI recommends Astra for the hardest science work.

For latency-sensitive calls that used none, start at low and measure. The reasoning guide describes low as efficient reasoning “with a modest latency increase.” To change effort mid-conversation without breaking the prompt cache, append a configuration_update input item rather than changing the request-level reasoning.effort.

Migrate from gpt-6-sol: four code changes

  1. Swap the model ID. Replace gpt-6-sol with gpt-6.1-sol, and keep it in config or an environment variable so rollback is one edit.
  2. Remap none and minimal. OpenAI’s guidance: use low instead of none, and start minimal at low and compare on representative tasks. On GPT-6 Astra, which also lacks none, sending it returns HTTP 400, so fix this before you move traffic.
  3. Strip sampling parameters. When effort isn’t none, remove temperature, top_p and top_logprobs (and logprobs in Chat Completions). Code that paired temperature with none on GPT-6 Sol needs this.
  4. Move Chat Completions tool calls to Responses. GPT-6 Sol allowed function calling in Chat Completions only at reasoning_effort: "none". That combination has no equivalent on 6.1 Sol.

Then re-run anything that depends on recency: the cutoff moves from April 20 to April 30, 2026. If you came to Sol from Astra, the Astra-to-Sol migration guide covers that earlier step.

Batch, Flex, Fast and cached-input pricing

Every tier keeps GPT-6 Sol’s shape, with the cached-input column halved. Prices per 1M tokens are from the API pricing page. The model page adds that a prompt over 272K input tokens is billed at 2x input and cache rates and 1.5x output for the full request, the same rule GPT-6 Sol uses:

Tier Input Cached input Cache writes Output
Standard $2.00 $0.10 $2.50 $10.00
Batch $1.00 $0.05 $1.25 $5.00
Flex $1.00 $0.05 $1.25 $5.00
Fast $4.00 $0.20 $5.00 $20.00
Standard, prompt over 272K input tokens $4.00 $0.20 $5.00 $15.00

Flex is a per-request service_tier: "flex". Fast is service_tier: "fast", with "priority" accepted as an alias. Fast mode isn’t available with EU data residency. Ultrafast for GPT-6.1 Sol is “coming soon” and is broadly available only for GPT-6 Astra today; see OpenAI Ultrafast mode. For overnight jobs, the OpenAI Batch API guide walks through a batch run.

The cache is where the upgrade saves money. Reads cost 0.05x the input rate on 6.1 Sol versus 0.1x on GPT-6 Sol, and writes cost 1.25x on both, per the prompt caching guide. Take a 50,000-token system prompt reused across 1,000 requests. One write costs $0.125 on either model; the 999 reads cost $9.99 on GPT-6 Sol and $5.00 on GPT-6.1 Sol. The minimum cacheable prefix is 1,024 visible tokens, and a cached prefix stays eligible for at least 30 minutes after its last write or reuse. For breakpoint strategy, see GPT-6 prompt caching.

Test the swap in Apidog

Don’t switch production on list prices alone. Send the same saved request to both IDs and compare what comes back. In Apidog:

  1. Create an environment with OPENAI_API_KEY (stored as a secret), MODEL_ID set to gpt-6-sol, and EFFORT set to medium.
  2. Create POST https://api.openai.com/v1/responses with the header Authorization: Bearer {{OPENAI_API_KEY}} and this body, then save it:
{
  "model": "{{MODEL_ID}}",
  "reasoning": {"effort": "{{EFFORT}}"},
  "max_output_tokens": 25000,
  "input": "Return a JSON object with keys risk and fix for this policy: retry any 5xx three times with no idempotency key."
}
  1. Add assertions: HTTP 200, $.status equals completed, $.output[*].type contains message, $.usage.output_tokens is greater than 0, and $.usage.output_tokens_details.reasoning_tokens exists. Then check the output shape your code depends on, such as valid JSON with the keys you parse.
  2. Add a post-processor script that turns usage into dollars, using the input split from OpenAI’s prompt caching guide:
const u = pm.response.json().usage;
const d = u.input_tokens_details || {};
const cached = d.cached_tokens || 0;
const writes = d.cache_write_tokens || 0;
const model = pm.environment.get("MODEL_ID");
const cachedRate = model === "gpt-6.1-sol" ? 0.10 : 0.20;
const cost = ((u.input_tokens - cached - writes) * 2 + cached * cachedRate
  + writes * 2.5 + u.output_tokens * 10) / 1e6;
console.log(model, "cost per call $", cost.toFixed(5));
  1. Send it, set MODEL_ID to gpt-6.1-sol, and send again. Compare reasoning_tokens, output_tokens, the answer and the logged cost. If you’re remapping from none, run the baseline at none and the candidate at low.

Then move the request and a handful of real prompts into a test scenario and run the pair from the Apidog CLI in CI. --env-var overrides a variable for one run, so a single scenario covers both models:

npm install -g apidog-cli
apidog run --access-token "$APIDOG_ACCESS_TOKEN" -t "$SCENARIO_ID" -e "$ENV_ID" \
  --env-var "MODEL_ID=gpt-6-sol" -r cli,junit
apidog run --access-token "$APIDOG_ACCESS_TOKEN" -t "$SCENARIO_ID" -e "$ENV_ID" \
  --env-var "MODEL_ID=gpt-6.1-sol" -r cli,junit

A failed assertion fails the job, and the JUnit reports give you both runs side by side. For assertions on outputs that vary run to run, see testing non-deterministic AI agents.

FAQ

Is GPT-6.1 Sol more expensive than GPT-6 Sol? No. Both list at $2 input and $10 output per 1M tokens. GPT-6.1 Sol’s cached input is $0.10 versus $0.20, so cache-heavy workloads get cheaper.

What should I do with reasoning.effort: "none"? GPT-6.1 Sol supports neither none nor minimal. Map both to low, remove temperature and top_p, and re-run your evals before switching.

Can I use GPT-6.1 Sol with Chat Completions? Yes, for requests without tools. Tool calling needs the Responses API.

Is there a free GPT-6.1 Sol API tier? No. API calls are billed per token from the first request. Is GPT-6.1 Sol free? covers the cheapest routes.

Next step

Save the first request, run it on gpt-6-sol at your current effort, then on gpt-6.1-sol, and compare usage and output on a prompt from your own traffic. Download Apidog to keep both runs as assertions you can rerun in CI. Weighing Anthropic instead? See GPT-6.1 Sol vs Claude Sonnet 5.5.

Explore more

What Is GPT-6.1 Sol?

What Is GPT-6.1 Sol?

GPT-6.1 Sol explained: model ID gpt-6.1-sol, $2/$10 pricing with $0.10 cached input, 922K max input, effort levels, and OpenAI's benchmarks vs Astra.

30 September 2026

OpenAI DevDay 2026 for API developers: what shipped, what it costs, and what's still coming soon

OpenAI DevDay 2026 for API developers: what shipped, what it costs, and what's still coming soon

OpenAI DevDay 2026 for API developers: GPT-6.1 Sol at $2/$10, Ultrafast on Astra, Agents API computer use, MCP Events, and what to change this week.

30 September 2026

What Is Claude Sonnet 5.5?

What Is Claude Sonnet 5.5?

What is Claude Sonnet 5.5? Anthropic's Sept 28, 2026 model: $2/$10 pricing, 1M context, benchmarks, what changed from Sonnet 5, and where to use it.

29 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

How to use GPT-6.1 Sol APl ?