To use the OpenAI Decisions API, send a POST request to https://api.openai.com/v1/decisions with "model": "gpt-6-luna", an input (text, images, or both), and a questions array where each question is a predicate, a choice, or a score. You get back typed answers with probabilities instead of text to parse, and you pay $0.10 per 1M input tokens, with no output, cache-read or cache-write charges. The endpoint is in public beta as of October 6, 2026.
This guide covers getting a key, the first call in curl, Python and JavaScript, reading each answer type, three questions on one support ticket, image input, thresholds, and a testing setup in Apidog. For when to pick the endpoint at all, start with what is the OpenAI Decisions API.
Decisions API request at a glance
| Field | What it takes |
|---|---|
model |
gpt-6-luna (the only model available in beta) |
input |
A string, or an array of user messages whose content is a string or parts of type input_text and input_image |
questions[].type |
predicate, choice, or score |
questions[].instructions |
Required; the question in plain words |
questions[].name |
Optional; echoed back in the answer (null if omitted) |
questions[].choices |
choice only; 2 to 255 unique {value, description} objects, value a string or boolean |
questions[].levels |
score only; ordered {label, description} objects, lowest first, indices from 0 |
safety_identifier |
Optional opaque end-user id, up to 128 chars |
Source: the Decisions API reference. There’s no temperature, stream, tools or text.format on this endpoint.
Get a key and make the first call
Create a key in the OpenAI dashboard (the OpenAI API key guide covers it), export it as OPENAI_API_KEY, and never paste it into code. Then ask one yes/no question:
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "The box arrived crushed and the screen is cracked.",
"questions": [
{"type": "predicate", "name": "damaged",
"instructions": "Does the customer report a damaged item?"}
]
}'
The response has three top-level fields: model, answers, and usage. This is the shape from OpenAI’s reference, with output_tokens at 0 because the endpoint bills no output:
{
"model": "gpt-6-luna",
"answers": [
{"type": "predicate", "name": "damaged", "probability": 0.95}
],
"usage": {
"input_tokens": 42,
"input_tokens_details": {"cached_tokens": 0, "cache_write_tokens": 0},
"output_tokens": 0,
"output_tokens_details": {"reasoning_tokens": 0},
"total_tokens": 42
}
}
The same call in Python (SDK 3.26.0 or later):
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
decision = client.decisions.create(
model="gpt-6-luna",
input="The box arrived crushed and the screen is cracked.",
questions=[
{"type": "predicate", "name": "damaged",
"instructions": "Does the customer report a damaged item?"}
],
)
print(decision.answers[0].probability)
And in JavaScript (SDK 7.30.0 or later):
import OpenAI from "openai";
const client = new OpenAI();
const decision = await client.decisions.create({
model: "gpt-6-luna",
input: "The box arrived crushed and the screen is cracked.",
questions: [
{ type: "predicate", name: "damaged",
instructions: "Does the customer report a damaged item?" },
],
});
console.log(decision.answers[0].probability);
Read the answer by type
Answers come back in the order you asked, each with a type. Switch on the type, because any question can come back as a refusal.
predicatereturnsprobability, an estimate from 0 to 1 that the condition is true.choicereturnschoice(the winningvalue),probabilitiesas an array of{value, probability}objects, andconfidence.scorereturnsscore,probabilitiesas an array of{value, label, probability}objects (wherevalueis the 0-based level index), andconfidence. The score is the probability-weighted average of the level indices, so it can land between levels: 1.1 means “between level 1 and level 2, close to 1”.refusalreturns onlytypeandname. The model declined that question; other questions in the same request can still receive answers.
for a in decision.answers:
if a.type == "refusal":
send_to_review(a.name)
elif a.type == "predicate":
flag = a.probability > 0.9
elif a.type == "choice":
route = a.choice if a.confidence > 0.8 else "review"
elif a.type == "score":
priority = round(a.score)
OpenAI’s guide draws the line this way: choice for categories without an order, such as departments; score for ordered levels, such as severity.
Three questions on one support ticket
Independent questions share one request and one input, and each question can use a different type. Here’s a predicate, a choice and a score on a single ticket:
{
"model": "gpt-6-luna",
"input": "I was charged twice for my order.",
"questions": [
{"type": "predicate", "name": "refund_requested",
"instructions": "Is the customer asking for money back?"},
{"type": "choice", "name": "department",
"instructions": "Which team should handle this ticket?",
"choices": [
{"value": "billing", "description": "Charges, refunds, invoices"},
{"value": "technical", "description": "Bugs and errors in the product"},
{"value": "shipping", "description": "Delivery and tracking"},
{"value": "other", "description": "Anything else"}
]},
{"type": "score", "name": "urgency",
"instructions": "How urgent is this ticket?",
"levels": [
{"label": "low", "description": "No time pressure"},
{"label": "medium", "description": "Needs a reply this week"},
{"label": "high", "description": "Customer is blocked or losing money"}
]}
]
}
The answers array comes back in the same order. The choice values below are OpenAI’s guide values for this exact input; the predicate and score values are illustrative:
"answers": [
{"type": "predicate", "name": "refund_requested", "probability": 0.88},
{"type": "choice", "name": "department", "choice": "billing",
"probabilities": [
{"value": "billing", "probability": 0.95},
{"value": "technical", "probability": 0.02},
{"value": "shipping", "probability": 0.01},
{"value": "other", "probability": 0.02}
],
"confidence": 0.93},
{"type": "score", "name": "urgency", "score": 1.6,
"probabilities": [
{"value": 0, "label": "low", "probability": 0.05},
{"value": 1, "label": "medium", "probability": 0.30},
{"value": 2, "label": "high", "probability": 0.65}
],
"confidence": 0.65}
]
Two rules from the guide: include a fallback such as other when your categories don’t cover every input, and write questions around observable criteria so adjacent score levels mean different things. If a second decision depends on the first answer, send a separate request.
Image input
Pass an image as a content part inside a user message. The guide documents inline base64 data URLs:
{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
{"type": "input_text", "text": "Photo attached to a return request."},
{"type": "input_image", "image_url": "data:image/jpeg;base64,/9j/4AAQ..."}
]
}],
"questions": [
{"type": "predicate", "name": "visible_damage",
"instructions": "Is the product visibly damaged?"}
]
}
The API reference also lists publicly accessible HTTP(S) URLs, up to 128 images across all messages in one request, and an optional detail field (low, high, auto, original), so test hosted URLs against your own account before relying on them. file_id inputs aren’t supported on either page.
Pick thresholds from labeled examples
OpenAI publishes no accuracy or calibration figures for the endpoint. Its guidance is to use labeled examples from your own application to set thresholds for routing, filtering or review, based on the cost of false positives versus false negatives. In practice that means a small CSV of real tickets with the department a human picked, run through the same request, so you can see where confidence separates clean routes from the ones that need a person. The next section builds that loop.
Test the Decisions API in Apidog
Saved requests make threshold tuning and regression checks repeatable. Here’s the setup in Apidog:
- Store the key as an environment variable. Create an environment, add
OPENAI_API_KEYas a secret variable (Apidog environments and secret variables shows the setup), and set theAuthorizationheader toBearer {{OPENAI_API_KEY}}. The key never lands in a shared request body. - Save one request per question type. Create a POST to
https://api.openai.com/v1/decisionswithContent-Type: application/json, paste thechoicequestion from the ticket example above on its own, and save it. Duplicate it for the predicate and score versions. - Add JSONPath assertions. On the choice request: status is 200,
$.answers[0].typeequalschoice,$.answers[0].choiceequalsbilling,$.answers[0].confidenceis greater than 0.8, and$.usage.output_tokensequals 0. For the damage predicate, assert$.answers[?(@.name=='damaged')].probabilityis greater than 0.9. A wording change in your instructions, or a model behaviour change, now fails a test instead of misrouting tickets. - Run it over labeled tickets. Build a test scenario from the saved request and attach a small CSV with two columns,
ticket_textandexpected_department. Map{{ticket_text}}intoinputand assert$.answers[0].choiceequals{{expected_department}}. The run report showsconfidencefor every row, which is the data OpenAI tells you to set thresholds from. The point below which every mis-route sits becomes your “route automatically” threshold in code. - Mock the
answersarray for the frontend. Point the router or UI at a mock of the same endpoint that returns achoiceanswer withconfidenceabove and below your threshold, plus arefusal, so the review-queue path gets built before you spend an input token. Conditional mock responses in Apidog covers switching mocks on request content. - Run the scenario in CI. Export an access token, then add a step to your pipeline:
apidog run --access-token "$APIDOG_ACCESS_TOKEN" \
-t "$SCENARIO_ID" -e "$ENV_ID" -r cli,junit
A failed assertion fails the build, so a quiet drop in confidence is caught before deploy rather than in the support queue. For broader patterns, see testing LLM applications.
Handle errors and edge cases
- 429 rate limited. No Decisions-specific limits are published; check Settings > Organization > Limits for your numbers. Back off exponentially and honor a
Retry-Afterheader when present. The rate limit exceeded guide has a retry wrapper. - Refusal answers. Treat
type: "refusal"as a routing outcome, not an exception. Send that ticket to a human and keep the other answers from the same request. - Dependent decisions. Every question is evaluated independently against the shared input. Anything that depends on an earlier answer needs its own request.
FAQ
How much does the Decisions API cost? $0.10 per 1M input tokens on gpt-6-luna, with no output, cache-read or cache-write charges. A 500-token ticket with three questions costs 500 / 1,000,000 x $0.10 = $0.00005, so a million such tickets cost $50. Long-context input over 272K tokens is 2x, and regional processing adds 10%.
Is the Decisions API free? No. There’s no free Decisions tier. If you want to try GPT-6 Luna without paying, the GPT-6 Luna free routes post lists what exists.
How fast is it? OpenAI says about 10x faster than the Responses API and publishes no absolute latency number. One developer on the OpenAI forum reported image decisions in about 0.8 seconds.
Which models work with the Decisions API? Only gpt-6-luna today. It’s an endpoint on Luna, not a separate model. See what is GPT-6 Luna for the model itself.
When should I use Structured Outputs instead? When you need an object in your own JSON schema, such as extracted fields or a written explanation, or function calling when the model should request a tool with arguments. The Decisions API vs Responses API post shows the same ticket done both ways.
How does it compare with Jev? Both return typed answers with probabilities and bill input only; Jev is text-only at $0.042 per 1M. The Decisions API vs Jev comparison has the full table.
Next step
Send the three-question ticket request from this guide, then run it over 20 of your own labeled tickets and see where confidence separates correct routes from wrong ones. Then download Apidog to keep the request, the CSV scenario and the assertions together, so the threshold you pick today gets re-checked on every deploy.



