How to Use the OpenAI Decisions API ?

How to use the OpenAI Decisions API: first call in curl, Python and JavaScript, predicate, choice and score answers, image input, and Apidog tests.

Ashley Innocent

Ashley Innocent

10 October 2026

How to Use the OpenAI Decisions API ?

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

To use the OpenAI Decisions API, send a POST request to https://api.openai.com/v1/decisions with "model": "gpt-6-luna", an input (text, images, or both), and a questions array where each question is a predicate, a choice, or a score. You get back typed answers with probabilities instead of text to parse, and you pay $0.10 per 1M input tokens, with no output, cache-read or cache-write charges. The endpoint is in public beta as of October 6, 2026.

This guide covers getting a key, the first call in curl, Python and JavaScript, reading each answer type, three questions on one support ticket, image input, thresholds, and a testing setup in Apidog. For when to pick the endpoint at all, start with what is the OpenAI Decisions API.

button

Decisions API request at a glance

Field What it takes
model gpt-6-luna (the only model available in beta)
input A string, or an array of user messages whose content is a string or parts of type input_text and input_image
questions[].type predicate, choice, or score
questions[].instructions Required; the question in plain words
questions[].name Optional; echoed back in the answer (null if omitted)
questions[].choices choice only; 2 to 255 unique {value, description} objects, value a string or boolean
questions[].levels score only; ordered {label, description} objects, lowest first, indices from 0
safety_identifier Optional opaque end-user id, up to 128 chars

Source: the Decisions API reference. There’s no temperature, stream, tools or text.format on this endpoint.

Get a key and make the first call

Create a key in the OpenAI dashboard (the OpenAI API key guide covers it), export it as OPENAI_API_KEY, and never paste it into code. Then ask one yes/no question:

curl https://api.openai.com/v1/decisions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "input": "The box arrived crushed and the screen is cracked.",
    "questions": [
      {"type": "predicate", "name": "damaged",
       "instructions": "Does the customer report a damaged item?"}
    ]
  }'

The response has three top-level fields: model, answers, and usage. This is the shape from OpenAI’s reference, with output_tokens at 0 because the endpoint bills no output:

{
  "model": "gpt-6-luna",
  "answers": [
    {"type": "predicate", "name": "damaged", "probability": 0.95}
  ],
  "usage": {
    "input_tokens": 42,
    "input_tokens_details": {"cached_tokens": 0, "cache_write_tokens": 0},
    "output_tokens": 0,
    "output_tokens_details": {"reasoning_tokens": 0},
    "total_tokens": 42
  }
}

The same call in Python (SDK 3.26.0 or later):

from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from the environment

decision = client.decisions.create(
    model="gpt-6-luna",
    input="The box arrived crushed and the screen is cracked.",
    questions=[
        {"type": "predicate", "name": "damaged",
         "instructions": "Does the customer report a damaged item?"}
    ],
)
print(decision.answers[0].probability)

And in JavaScript (SDK 7.30.0 or later):

import OpenAI from "openai";

const client = new OpenAI();

const decision = await client.decisions.create({
  model: "gpt-6-luna",
  input: "The box arrived crushed and the screen is cracked.",
  questions: [
    { type: "predicate", name: "damaged",
      instructions: "Does the customer report a damaged item?" },
  ],
});
console.log(decision.answers[0].probability);

Read the answer by type

Answers come back in the order you asked, each with a type. Switch on the type, because any question can come back as a refusal.

for a in decision.answers:
    if a.type == "refusal":
        send_to_review(a.name)
    elif a.type == "predicate":
        flag = a.probability > 0.9
    elif a.type == "choice":
        route = a.choice if a.confidence > 0.8 else "review"
    elif a.type == "score":
        priority = round(a.score)

OpenAI’s guide draws the line this way: choice for categories without an order, such as departments; score for ordered levels, such as severity.

Three questions on one support ticket

Independent questions share one request and one input, and each question can use a different type. Here’s a predicate, a choice and a score on a single ticket:

{
  "model": "gpt-6-luna",
  "input": "I was charged twice for my order.",
  "questions": [
    {"type": "predicate", "name": "refund_requested",
     "instructions": "Is the customer asking for money back?"},
    {"type": "choice", "name": "department",
     "instructions": "Which team should handle this ticket?",
     "choices": [
       {"value": "billing", "description": "Charges, refunds, invoices"},
       {"value": "technical", "description": "Bugs and errors in the product"},
       {"value": "shipping", "description": "Delivery and tracking"},
       {"value": "other", "description": "Anything else"}
     ]},
    {"type": "score", "name": "urgency",
     "instructions": "How urgent is this ticket?",
     "levels": [
       {"label": "low", "description": "No time pressure"},
       {"label": "medium", "description": "Needs a reply this week"},
       {"label": "high", "description": "Customer is blocked or losing money"}
     ]}
  ]
}

The answers array comes back in the same order. The choice values below are OpenAI’s guide values for this exact input; the predicate and score values are illustrative:

"answers": [
  {"type": "predicate", "name": "refund_requested", "probability": 0.88},
  {"type": "choice", "name": "department", "choice": "billing",
   "probabilities": [
     {"value": "billing", "probability": 0.95},
     {"value": "technical", "probability": 0.02},
     {"value": "shipping", "probability": 0.01},
     {"value": "other", "probability": 0.02}
   ],
   "confidence": 0.93},
  {"type": "score", "name": "urgency", "score": 1.6,
   "probabilities": [
     {"value": 0, "label": "low", "probability": 0.05},
     {"value": 1, "label": "medium", "probability": 0.30},
     {"value": 2, "label": "high", "probability": 0.65}
   ],
   "confidence": 0.65}
]

Two rules from the guide: include a fallback such as other when your categories don’t cover every input, and write questions around observable criteria so adjacent score levels mean different things. If a second decision depends on the first answer, send a separate request.

Image input

Pass an image as a content part inside a user message. The guide documents inline base64 data URLs:

{
  "model": "gpt-6-luna",
  "input": [{
    "role": "user",
    "content": [
      {"type": "input_text", "text": "Photo attached to a return request."},
      {"type": "input_image", "image_url": "data:image/jpeg;base64,/9j/4AAQ..."}
    ]
  }],
  "questions": [
    {"type": "predicate", "name": "visible_damage",
     "instructions": "Is the product visibly damaged?"}
  ]
}

The API reference also lists publicly accessible HTTP(S) URLs, up to 128 images across all messages in one request, and an optional detail field (low, high, auto, original), so test hosted URLs against your own account before relying on them. file_id inputs aren’t supported on either page.

Pick thresholds from labeled examples

OpenAI publishes no accuracy or calibration figures for the endpoint. Its guidance is to use labeled examples from your own application to set thresholds for routing, filtering or review, based on the cost of false positives versus false negatives. In practice that means a small CSV of real tickets with the department a human picked, run through the same request, so you can see where confidence separates clean routes from the ones that need a person. The next section builds that loop.

Test the Decisions API in Apidog

Saved requests make threshold tuning and regression checks repeatable. Here’s the setup in Apidog:

  1. Store the key as an environment variable. Create an environment, add OPENAI_API_KEY as a secret variable (Apidog environments and secret variables shows the setup), and set the Authorization header to Bearer {{OPENAI_API_KEY}}. The key never lands in a shared request body.
  2. Save one request per question type. Create a POST to https://api.openai.com/v1/decisions with Content-Type: application/json, paste the choice question from the ticket example above on its own, and save it. Duplicate it for the predicate and score versions.
  3. Add JSONPath assertions. On the choice request: status is 200, $.answers[0].type equals choice, $.answers[0].choice equals billing, $.answers[0].confidence is greater than 0.8, and $.usage.output_tokens equals 0. For the damage predicate, assert $.answers[?(@.name=='damaged')].probability is greater than 0.9. A wording change in your instructions, or a model behaviour change, now fails a test instead of misrouting tickets.
  4. Run it over labeled tickets. Build a test scenario from the saved request and attach a small CSV with two columns, ticket_text and expected_department. Map {{ticket_text}} into input and assert $.answers[0].choice equals {{expected_department}}. The run report shows confidence for every row, which is the data OpenAI tells you to set thresholds from. The point below which every mis-route sits becomes your “route automatically” threshold in code.
  5. Mock the answers array for the frontend. Point the router or UI at a mock of the same endpoint that returns a choice answer with confidence above and below your threshold, plus a refusal, so the review-queue path gets built before you spend an input token. Conditional mock responses in Apidog covers switching mocks on request content.
  6. Run the scenario in CI. Export an access token, then add a step to your pipeline:
apidog run --access-token "$APIDOG_ACCESS_TOKEN" \
  -t "$SCENARIO_ID" -e "$ENV_ID" -r cli,junit

A failed assertion fails the build, so a quiet drop in confidence is caught before deploy rather than in the support queue. For broader patterns, see testing LLM applications.

Handle errors and edge cases

FAQ

How much does the Decisions API cost? $0.10 per 1M input tokens on gpt-6-luna, with no output, cache-read or cache-write charges. A 500-token ticket with three questions costs 500 / 1,000,000 x $0.10 = $0.00005, so a million such tickets cost $50. Long-context input over 272K tokens is 2x, and regional processing adds 10%.

Is the Decisions API free? No. There’s no free Decisions tier. If you want to try GPT-6 Luna without paying, the GPT-6 Luna free routes post lists what exists.

How fast is it? OpenAI says about 10x faster than the Responses API and publishes no absolute latency number. One developer on the OpenAI forum reported image decisions in about 0.8 seconds.

Which models work with the Decisions API? Only gpt-6-luna today. It’s an endpoint on Luna, not a separate model. See what is GPT-6 Luna for the model itself.

When should I use Structured Outputs instead? When you need an object in your own JSON schema, such as extracted fields or a written explanation, or function calling when the model should request a tool with arguments. The Decisions API vs Responses API post shows the same ticket done both ways.

How does it compare with Jev? Both return typed answers with probabilities and bill input only; Jev is text-only at $0.042 per 1M. The Decisions API vs Jev comparison has the full table.

Next step

Send the three-question ticket request from this guide, then run it over 20 of your own labeled tickets and see where confidence separates correct routes from wrong ones. Then download Apidog to keep the request, the CSV scenario and the assertions together, so the threshold you pick today gets re-checked on every deploy.

Explore more

Claude Haiku 5.5 vs GPT-6 Luna

Claude Haiku 5.5 vs GPT-6 Luna

Haiku 5.5 vs GPT-6 Luna: same $0.10/$0.50 price under 100K tokens, different long-prompt tiers, Anthropic's benchmarks, and per-effort cost data.

8 October 2026

Claude Haiku 5.5 Benchmarks

Claude Haiku 5.5 Benchmarks

Claude Haiku 5.5 benchmarks: 72.4% OSWorld, 1620 GDPval-AA, 39.2% Terminal-Bench at max effort. Who ran each test, per-effort costs, and your own eval.

8 October 2026

Claude Haiku 5.5 vs Haiku 4.5: What Changed and the Breaking Changes to Fix First

Claude Haiku 5.5 vs Haiku 4.5: What Changed and the Breaking Changes to Fix First

Haiku 5.5 vs Haiku 4.5: 90% cheaper up to 100K tokens, 1M context, and five breaking changes that return 400s. Before/after JSON fixes inside.

8 October 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

How to Use the OpenAI Decisions API ?