The OpenAI Decisions API is a POST /v1/decisions endpoint, running on GPT-6 Luna, that takes text or images plus a list of questions and returns typed answers instead of prose: a predicate probability, a choice with per-option probabilities, or a score over ordered levels. Input costs $0.10 per 1M tokens with no output, cache-read or cache-write charges, and the endpoint has been in public beta since 2026-10-06, with OpenAI saying GA is expected “in the coming weeks”.
This post covers what the endpoint returns, what it costs, where it fits next to Structured Outputs and function calling, and how to test it. For the walkthrough with curl, Python and JavaScript, read how to use the OpenAI Decisions API next; if you’re already on the Responses API, the Decisions vs Responses comparison shows the same job done both ways. Throughout, we’ll use Apidog to store the key, save requests, and assert on the answers array so a change in model behaviour fails a test instead of misrouting a ticket.
Anatomy of a Decisions request and response
Three request fields, three response fields. No id, no generated text, nothing to parse.
| Part | Field | What it holds |
|---|---|---|
| Request | model |
gpt-6-luna, the only model available today |
| Request | input |
A string, or an array of user messages whose content mixes input_text and input_image parts |
| Request | questions |
An array of questions, each with a type, required instructions, and an optional name |
| Request | safety_identifier |
Optional end-user id, up to 128 characters |
| Response | model |
Echoes gpt-6-luna |
| Response | answers |
One entry per question, in the order you asked, with type and name |
| Response | usage |
input_tokens, input_tokens_details, output_tokens, output_tokens_details, total_tokens |
Note what’s missing: no temperature, reasoning, stream, store, tools or text.format. For those you want the Responses API. And output_tokens is 0 in OpenAI’s own reference example, which is why the pricing below has no output line.
The three question types
Each question carries its own type, and you can mix types on one input. Put independent questions in the same request; for decisions that depend on an earlier answer, OpenAI’s guide says to send separate requests.
predicate: a yes/no probability
A predicate asks whether a condition holds and returns a probability from 0 to 1 that it’s true.
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "The box arrived crushed and the screen is cracked.",
"questions": [
{"type": "predicate", "name": "damaged",
"instructions": "Is the product described as damaged?"}
]
}'
OpenAI’s reference example for this shape returns:
{
"model": "gpt-6-luna",
"answers": [
{"type": "predicate", "name": "damaged", "probability": 0.95}
],
"usage": {
"input_tokens": 42,
"input_tokens_details": {"cached_tokens":0,"cache_write_tokens":0},
"output_tokens": 0,
"output_tokens_details": {"reasoning_tokens":0},
"total_tokens": 42
}
}
choice: one label from an unordered set
A choice adds a choices array of {value, description} objects: 2 to 255 unique choices, where value is a string or a boolean (true and "true" are distinct). OpenAI recommends a fallback such as other when your categories don’t cover every input.
{
"model": "gpt-6-luna",
"input": "I was charged twice for my order.",
"questions": [
{"type": "choice", "name": "department",
"instructions": "Which team should handle this ticket?",
"choices": [
{"value":"billing"}, {"value":"technical"},
{"value":"shipping"}, {"value":"other"}
]}
]
}
The guide’s illustrative answer for this input:
{"type": "choice", "name": "department", "choice": "billing",
"probabilities": [
{"value":"billing","probability":0.95},
{"value":"technical","probability":0.02},
{"value":"shipping","probability":0.01},
{"value":"other","probability":0.02}
],
"confidence": 0.93}
score: a position on an ordered scale
A score adds levels, an array of {label, description} ordered lowest to highest. Indices start at 0, and the returned score is the probability-weighted average of those indices, so it can land between levels.
{
"model": "gpt-6-luna",
"input": "Export fails in Safari but works in Chrome.",
"questions": [
{"type": "score", "name": "severity",
"instructions": "How badly does this bug block the user?",
"levels": [
{"label":"Cosmetic"},
{"label":"Workaround available"},
{"label":"Fully blocked"}
]}
]
}
In the guide’s example the probabilities are 0.1, 0.7 and 0.2 across the three levels, giving a score of 1.1 and a confidence of 0.55. Read 1.1 as “between level 1 and level 2, close to 1”. The guide’s rule: choice for unordered categories such as departments; score for ordered levels such as severity.
A fourth answer type, refusal, can appear for any single question as {"type":"refusal","name":...}. Other questions in the same request can still get answers, so branch on type before reading a field.
Speed, as OpenAI describes it
OpenAI says the Decisions API is about 10x faster than the Responses API; the announcement phrases it as up to 10x faster than GPT-6 Luna through Responses. OpenAI publishes no absolute latency number. One developer on the OpenAI forum reported image-input decisions returning in about 0.8 seconds on a slow connection: an anecdote, not a benchmark. Measure your own p95 before promising anything.
Pricing: $0.10 per million input tokens, nothing else
With gpt-6-luna, input costs $0.10 per 1M tokens. You pay only for input tokens: there are no cache-read, cache-write, or output-token charges. The usage object carries cached_tokens and cache_write_tokens fields, but per a reply on OpenAI’s developer forum there’s no caching on Decisions yet, so expect 0.
Two multipliers apply. Input over 272K tokens is billed at 2x, which works out to $0.20 per 1M (derived from the pricing page’s long-context multiplier). Regional processing through the US or EU data-residency endpoints adds 10%. No Batch, Flex or Fast tier is documented for /v1/decisions, so don’t plan around a discount that exists only on Responses.
Here’s the arithmetic for a support-routing workload. A 500-token ticket with three questions in one request costs 500 / 1,000,000 x $0.10 = $0.00005. One million such tickets cost $50. The same ticket through the Responses API with a 40-token JSON label at $0.50 per 1M output adds 40 / 1,000,000 x $0.50 = $0.00002 per request on top of the input, before reasoning tokens, which Luna bills as output on Responses and Decisions doesn’t bill at all. The honest framing is “Decisions bills no output tokens”, not a percentage. For the full Luna rate card and what prompt caching does on Responses, see what is GPT-6 Luna.
When to use Decisions, Structured Outputs, or function calling
OpenAI draws the line itself: use Structured Outputs with the Responses API when you need an object that follows your own JSON schema, such as extracted fields or a written explanation, or function calling when you need the model to request a tool call with arguments. Decisions is for classifying content, routing requests, and prioritizing work.
| You need | Use |
|---|---|
| A label, a probability, or a severity with confidence | Decisions API |
| An object in your own JSON schema (extracted fields, an explanation) | Structured Outputs on Responses |
| The model to pick a tool and fill its arguments | Function calling on Responses |
| Streaming, conversation state, tools, caching, or Batch | Responses API |
A Structured Outputs enum can return a label. It can’t return a probability distribution or a confidence field unless you ask the model to write one, and then it’s generated text, not a measured probability. Decisions gives you numbers you can threshold. OpenAI tells you to set those thresholds from labeled examples in your own application, weighing the cost of false positives against false negatives, because no accuracy or calibration figures are published. Weighing a second typed-decision vendor? The Decisions vs Jev comparison covers price, inputs and output shapes side by side.
Images, and the base64 caveat
input accepts user messages whose content mixes input_text and input_image parts, with an optional detail of low, high, auto (the default) or original. The guide says images must be inline base64 data URLs; hosted URLs and file_id aren’t supported. The API reference also lists publicly accessible HTTP(S) URLs, up to 128 images per request. Treat base64 as the documented path and test a hosted URL before relying on it.
Data controls
The Decisions API supports Zero Data Retention and HIPAA use for eligible customers. Data residency and regional processing are supported in the United States and Europe (EEA plus Switzerland) via us.api.openai.com and eu.api.openai.com. The endpoint is reachable from every supported API region, though availability in a region doesn’t imply inference runs there. Abuse-monitoring logs are retained up to 30 days by default. If you’re routing patient messages, read our HIPAA API compliance guide first.
Availability: beta now, GA soon
The endpoint went to public beta for all developers on 2026-10-06 and sits under “Beta APIs” in the reference. OpenAI’s guide says GA is expected “in the coming weeks”; no date is given. SDK examples require Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0 or Java 4.78.0 or later; the call is client.decisions.create(...) in Python and JavaScript. A Playground at platform.openai.com/decisions lets you try questions before writing code. No Decisions-specific rate limits are published; check your organization’s limits page. There’s no free Decisions tier; for no-cost Luna access, see our Luna free routes post.
Testing Decisions calls in Apidog
Typed answers are easy to assert on, which is the point. Three steps cover most teams.
Store the key once. Put OPENAI_API_KEY in an Apidog environment variable and reference {{OPENAI_API_KEY}} in the Authorization: Bearer header, so the literal key never lands in a shared request.
Save one request per question type, with JSONPath assertions: status 200, $.answers[0].type equals choice, $.answers[0].choice equals billing, $.answers[0].confidence greater than 0.8, $.answers[?(@.name=='damaged')].probability greater than 0.9, and $.usage.output_tokens equals 0, which catches a billing surprise before your invoice does.
Pick thresholds from a labeled set. Build a test scenario in Apidog that runs the same request over a CSV of ticket text and expected department, then set the auto-route threshold where the false-positive cost crosses the review-queue cost. Run it in CI with the Apidog CLI so a model or alias change fails a test instead of a customer. The how-to guide covers each step, including mocking the answers array so the frontend can be built before the router is final.
FAQ
Is the Decisions API a new model? No. It’s an endpoint, POST /v1/decisions, that runs on GPT-6 Luna. Luna launched 2026-09-22; the endpoint went to public beta on 2026-10-06.
What does the Decisions API cost? $0.10 per 1M input tokens with no output, cache-read or cache-write charges. Input over 272K tokens is 2x, and regional processing adds 10%.
Does it return my own JSON schema? No. It returns answers with probability, choice, or score fields. For your own schema, use Structured Outputs on the Responses API.
How accurate is it? OpenAI publishes no accuracy or calibration figures. Set thresholds from your own labeled data; a data-driven LLM test scenario is the practical way.
Where to start
Pick one routing decision your app makes today with a regex or a prompt-and-parse loop, write it as a single choice question with an other fallback, and run it over 50 labeled examples. If the confidence distribution separates cleanly, you have a threshold and a test. If it doesn’t, the question needs sharper criteria. To run that experiment with saved requests and assertions, download Apidog and import the curl above.



