Use the Decisions API when the job is to classify, route, score or gate something and you want probabilities back: it runs on GPT-6 Luna, returns typed answers instead of text, bills input only at $0.10 per 1M tokens with no output, cache-read or cache-write charges, and OpenAI says it’s about 10x faster than the Responses API. Use the Responses API when you need generated text, JSON in your own schema, tool calls, streaming or conversation state. Decisions entered public beta on 2026-10-06.
This post runs one job (routing a support ticket) through both endpoints, compares what each returns, works the cost once, and closes with a migration note and a way to test both in one Apidog project. For the endpoint’s anatomy start with the Decisions API pillar; for the baseline, see our Responses API guide.
Feature matrix
| Decisions API | Responses API (GPT-6 Luna) | |
|---|---|---|
| Endpoint | POST /v1/decisions |
POST /v1/responses |
| Output | predicate, choice, score answers (plus refusal) with probabilities and confidence from the endpoint |
Generated text, or JSON that follows your schema via text.format |
| Your own JSON schema | No | Yes, json_schema with strict: true |
| Tools / function calling | No | Yes |
| Streaming | No | Yes |
| Conversation state | No | Yes |
| Prompt caching | No cache charges; per OpenAI’s forum, no caching yet | Yes, cached input $0.01 per 1M |
| Batch | Not documented | Yes, 50% of standard |
| Images | Yes, base64 data URLs; the reference also lists public HTTP(S) URLs, up to 128 per request | Yes, Luna takes text and images |
| Chained (dependent) decisions | Separate requests | One generated response can carry dependent fields |
| Price per 1M, short context | $0.10 input; no output charge | $0.10 input, $0.50 output including reasoning tokens |
| ZDR / HIPAA | Supported for eligible customers; regional processing in the US and EU | Not covered in this comparison; see OpenAI’s data controls page |
Every row comes from OpenAI’s Decisions guide, the API reference and the pricing page.
The same job both ways: route a support ticket
The ticket reads “I was charged twice for my order.” The departments are billing, technical, shipping and other. Here is the Responses request with Structured Outputs, which is how most teams do this today:
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "Route this support ticket to one department.\n\nTicket: I was charged twice for my order.",
"text": {
"format": {
"type": "json_schema",
"name": "ticket_route",
"strict": true,
"schema": {
"type": "object",
"properties": {
"department": {
"type": "string",
"enum": ["billing", "technical", "shipping", "other"]
}
},
"required": ["department"],
"additionalProperties": false
}
}
}
}'
And the Decisions request for the same ticket:
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "I was charged twice for my order.",
"questions": [
{
"type": "choice",
"name": "department",
"instructions": "Which department should handle this ticket?",
"choices": [
{"value": "billing", "description": "Charges, refunds, invoices"},
{"value": "technical", "description": "Bugs, errors, login problems"},
{"value": "shipping", "description": "Delivery, tracking, returns in transit"},
{"value": "other", "description": "Anything else"}
]
}
]
}'
The Responses body carries the question inside the prompt and the allowed answers inside a schema. The Decisions body carries the raw ticket as input and the question as a choice with 2 to 255 unique values; it has no temperature, reasoning, stream or text fields, because none exist on that endpoint.
What each one returns
Responses returns generated text. With a strict schema that text is valid JSON, so after parsing you hold a label:
{"department": "billing"}
If you want a confidence number you add a field to the schema and ask the model to write one; what comes back is generated text that looks like a probability, not a measured one.
Decisions returns the label plus the distribution behind it. The numbers below are OpenAI’s guide example for this exact input:
{
"model": "gpt-6-luna",
"answers": [
{
"type": "choice",
"name": "department",
"choice": "billing",
"probabilities": [
{"value": "billing", "probability": 0.95},
{"value": "technical", "probability": 0.02},
{"value": "shipping", "probability": 0.01},
{"value": "other", "probability": 0.02}
],
"confidence": 0.93
}
]
}
A usage object follows answers (shown in the cost section). No parser, no regex. The confidence field is what you threshold, and OpenAI’s guidance is to set that threshold from your own labeled examples, because no accuracy or calibration figures are published. A refusal arrives as {"type": "refusal", "name": "department"}; other questions in the same request still get answers.
Cost: the arithmetic once
Both endpoints bill Luna input at $0.10 per 1M tokens in short context (up to 272K input tokens). The split is on output. Take a 500-token ticket at 1,000,000 requests:
- Decisions: 500 / 1,000,000 x $0.10 = $0.00005 per request, so $50 for the million, with no output or cache lines to add.
- Responses: the same $50 of input, plus output at $0.50 per 1M. A 40-token JSON label is 40 / 1,000,000 x $0.50 = $0.00002 per request, or $20 for the million. Then add reasoning tokens, which Luna bills as output at the same $0.50.
So the visible gap on the label alone is $50 against $70. The larger gap is the reasoning line, and the honest way to state it is that Decisions bills no output tokens at all; both counters read 0 in OpenAI’s reference example:
"usage": {
"input_tokens": 42,
"input_tokens_details": {"cached_tokens": 0, "cache_write_tokens": 0},
"output_tokens": 0,
"output_tokens_details": {"reasoning_tokens": 0},
"total_tokens": 42
}
Two caveats. Responses has levers Decisions doesn’t: reasoning.effort goes down to none on Luna, prompt caching drops repeated input to $0.01 per 1M, and the Batch API halves standard rates. None of those is documented for Decisions. And long-context input (over 272K tokens) doubles the input rate on both, so a long Decisions request is $0.20 per 1M input (derived from the pricing page’s multiplier); regional processing adds 10%.
Speed
OpenAI says the Decisions API is about 10x faster than the Responses API. No absolute latency number is published, so treat the claim as a direction rather than a budget and measure your own p50 and p95 before you move a hot path. One developer on the OpenAI forum reported image-input decisions returning in about 0.8 seconds on a slow connection; that’s an anecdote, not a benchmark. The direction is plausible: Responses generates tokens, reasoning included, and you wait for the last one.
The decision rule
Pick Decisions when the output is one of these:
- A yes/no with a probability (
predicate): “Is this message spam?” - One of N unordered categories (
choice): department, intent, which model or tool to call next. Include a fallback such as “other”. - An ordered level (
score): severity, priority, urgency. The score is a probability-weighted average of 0-based level indices, so 1.1 means between level 1 and level 2, close to 1. - A gate: compare
confidenceorprobabilityto a threshold and send low-confidence items to a human queue.
Pick Responses when any of these is true:
- You need text a person will read: a summary, a reply, an explanation.
- You need an object in your own shape: extracted fields, nested structures, arrays of unknown length. That is Structured Outputs territory, and OpenAI’s guide says so.
- The model should request a tool call with arguments: function calling.
- You need streaming, conversation state, or a model other than Luna.
- One decision depends on another and you want both in one round trip. Decisions holds several independent questions on one input, but dependent decisions need separate requests.
Many pipelines want both: Decisions to classify and gate, Responses to write the reply.
Migrating a classifier from Responses to Decisions
If you already route tickets with a strict enum schema, the move is small:
- Keep the same
input, stripped back to the raw ticket; the question moves out of the prompt. - Put the question in
questionsas achoice, with your enum values aschoices[].valueand a one-linedescriptioneach. Values can be strings or booleans, andtrueand"true"are distinct. - Delete the parser. Read
answers[0].choiceandanswers[0].confidence; answers arrive in the order you asked and echo thenameyou set. Then set a threshold from a labeled sample. - Check the input path. Decisions accepts user messages only: no system or assistant roles, no function calls, no files, no
file_id. Fold system-prompt rules intoinstructionsor the choice descriptions. Images go in as base64 data URLs; the reference also lists public HTTP(S) URLs, so test hosted images first. - Split chains. “Classify, then if billing decide refund eligibility” becomes two requests.
Test both in one Apidog project
The cleanest way to decide is to run both requests against the same labeled tickets and compare. In Apidog, store the key once as an environment variable and reference {{OPENAI_API_KEY}} in the Authorization: Bearer header of both saved requests, so no literal key lands in a saved body.
Give both requests the same assertion: the department equals billing. On the Decisions request that is a JSONPath assertion on $.answers[0].choice, with $.answers[0].confidence greater than 0.8 and $.usage.output_tokens equals 0 beside it. On the Responses request the label sits inside generated text, so a short post-request script parses it into a variable the assertion checks. Then compare usage on the two responses: Decisions reports zero output and reasoning tokens, Responses doesn’t.
Turn the pair into a data-driven test scenario over a CSV of ticket text and expected department, and the run shows how many tickets each endpoint routes correctly above your confidence line. Mock the answers array so the router can be built first, as in conditional mock responses, and run the scenario in CI with the Apidog CLI so a wording or model-alias change fails a test instead of misrouting tickets. See testing LLM applications for more assertion patterns.
FAQ
Can the Responses API return probabilities like Decisions does? Not as measured values. A confidence field in a JSON schema gets you a number the model wrote, which is generated text. Decisions returns probabilities over the options you supplied from the endpoint itself.
Can I use a model other than GPT-6 Luna on Decisions? No. The guide states gpt-6-luna is the only model currently available. See our GPT-6 Luna overview.
How is Decisions different from TypeSafe’s Jev? Both return typed answers with probabilities and bill input only; they differ on price, inputs and response shapes. See Decisions API vs Jev.
Is the Decisions API free? No. It bills $0.10 per 1M input tokens, with no free Decisions tier documented. For free routes to Luna itself, see how to use GPT-6 Luna for free.
Next step
Take one classifier you run through Responses today, rebuild it as a choice question, and run both over 50 labeled tickets in Apidog with the same assertion. If the confidence threshold holds and usage shows zero output tokens, you have your answer. Download Apidog, then follow how to use the Decisions API for the first call and the full testing walkthrough.



