OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers
OpenAI has released the Decisions API in public beta. It turns text and images into typed answers your code can branch on. OpenAI team states the OpenAI Decisions API runs about 10x faster than the Responses API. It targets a common pattern: prompt an LLM, then parse its text into a label.
TL;DR
- Size: GPT-6 Luna parameter count is not disclosed. Its model card lists a 1,050,000-token context window.
- Runs on: OpenAI-hosted API only, via
POST /v1/decisions. No open weights, no self-hosting. - Performance: About 10x faster than the Responses API, per OpenAI.
- Best: $0.10 per 1M input tokens, with no output, cache-read or cache-write charges.
- Bottom line:
- Best: fast, typed decisions with probabilities.
- Worst: 1 model, beta status, no independent evals yet.
- Best: fast, typed decisions with probabilities.
What is the OpenAI Decisions API?
The Decisions API is an OpenAI endpoint that evaluates text, images or both and returns typed answers. It does not generate prose. A request has 3 fields: model, input and questions. The only supported model today is gpt-6-luna. OpenAI expects general availability in the coming weeks.
How does the Decisions API work?
Each request carries shared evidence plus a list of questions. Each question has a unique name, a type and instructions. The response returns an answers array keyed by those names.
What are the 3 question types?
- predicate: checks a condition and returns a
probabilityfrom 0 to 1. Example: does a product photo show a crack, tear or dent? - choice: picks 1 value from options you supply. It also returns per-option
probabilitiesand aconfidencefield. - score: rates input against ordered
levels, indexed from 0. The score is a probability-weighted average of level indices.
OpenAI’s severity example makes the math concrete. Level probabilities of 0.1, 0.7 and 0.2 yield a score of 1.1. That value sits between ‘Workaround available’ and ‘Fully blocked’.
When should you use Structured Outputs instead?
OpenAI draws a clear line. Use Decisions for probabilities, choices or scores. Use Structured Outputs to fill your own JSON schema or write explanations. Use function calling when a model must request a tool call with arguments.
How fast is it, and what is the evidence?
OpenAI’s document claims about 10x faster responses than the Responses API. DevDay coverage put a decision near 150 ms, versus about 1.6 seconds for regular Luna calls. OpenAI has not published accuracy or calibration data for the endpoint. The docs advise setting thresholds with labeled examples from your own application.
How much does it cost, and where can you deploy it?
With gpt-6-luna, input costs $0.10 per 1M tokens. There are no output-token, cache-read or cache-write charges. Regional processing premiums and long-context multipliers still apply. Luna’s model card prices prompts above 272K tokens at 2x input rates.
- Endpoint:
POST /v1/decisionson the OpenAI API, plus a Playground. - SDK minimums: Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, Java 4.78.0.
- Compliance: Zero Data Retention and HIPAA for eligible customers.
- Residency: United States and Europe (EEA plus Switzerland).
- Voice: decisions can drive actions via client delegation with the Live API.
How does it compare with TypeSafe Jev?
The closest rival is TypeSafe Jev, launched September 15, 2026. Jev is a System One model that returns typed values with calibrated probabilities. TypeSafe prices input at $0.042 per 1M tokens, with free output. That makes OpenAI’s base rate about 2.4x higher. TypeSafe reports 70 to 500 ms end-to-end, measured from the US West Coast. Jev supports up to 255 choices but remains in early access. OpenAI’s edge is image input, compliance options and open beta access.
| Feature | OpenAI Decisions API | TypeSafe Jev 1.13 | GPT-6 Luna via Responses API |
|---|---|---|---|
| Release | Beta, Oct 6, 2026 | Early access, Sep 15, 2026 | GA, Sep 22, 2026 |
| Underlying model | gpt-6-luna | Jev (System One model) | gpt-6-luna |
| Parameters | Not disclosed | Not disclosed | Not disclosed |
| Input | Text, images (inline base64) | Text and structured state | Text, images |
| Output | Probability, choice or score | Typed values with probabilities | Generated text (JSON via Structured Outputs) |
| Max options per question | Not disclosed | Up to 255 | Not applicable |
| Speed (vendor-stated) | About 10x faster than Responses API | 70 to 500 ms end-to-end | Baseline |
| Input price per 1M tokens | $0.10 | $0.042 | $0.10 |
| Output price per 1M tokens | $0 | $0 | $0.50 |
| Compliance | ZDR, HIPAA (eligible); US and EU residency | Not disclosed | EU data residency available |
Key Takeaways
- OpenAI Decisions API returns probabilities, choices and scores, not prose.
- Input costs $0.10 per 1M tokens; output is free.
- OpenAI claims about 10x faster than the Responses API.
- TypeSafe Jev is cheaper at $0.042 per 1M input tokens.
Check out the Technical details here. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.