OpenAI's Decisions API Promises Routing Answers 10x Faster for Developers
OpenAI's Decisions API moves from limited preview to open beta, promising sub-200ms classification, routing, and scoring calls powered by GPT-6 Luna.
- OpenAI's Decisions API is now in public beta for all developers, powered by GPT-6 Luna.
- Returns decisions in roughly 150ms, about 10x faster than the same model via the Responses API.
- Supports three structured outputs: Predicates (probabilities), Choices (with confidence), and Scores (numeric ranges).
- Accepts text and image inputs; built for routing, classification, agent next-step selection, and safety gates.
- Early third-party tests show it competitive with Jev on accuracy, mixed on latency across tasks.
- Pricing, tuning support, and candidate-answer limits remain unpublished at beta launch.
OpenAI gives fast, bounded decisions their own API
OpenAI has moved its Decisions API from limited preview to a public beta available to all developers, according to its beta announcement. The endpoint packages a common application pattern involving prompts, JSON schemas, validation, and post-processing into a bounded decision call. OpenAI says requests can return up to 10 times faster than GPT-6 Luna requests made through the Responses API.
A Decisions request defines a predicate, choice set, or scoring range, then supplies context as text or images. The API returns a structured result that an application can use to route a request, classify content, evaluate an input, or choose an agent’s next action.
Bounded answers in three shapes
The Decisions API docs describe three output types:
| Type | What it returns | Typical use |
|---|---|---|
| Predicate | The estimated probability that a statement is true | Policy checks, condition gates, and risk detection |
| Choice | A selection from predefined options, with confidence scores | Classification, routing, and action selection |
| Score | An evaluation within a defined numeric range | Ranking, quality assessment, and rubric-based grading |
These structures cover boolean gates, multiclass classification, and numeric evaluation without requiring applications to parse generated prose or repair malformed JSON. Confidence values remain model estimates, so production thresholds need validation against representative labeled data and the costs of each error type.
The 10x claim needs a benchmark
OpenAI’s launch demo routed 10,000 customer requests among Billing, Technical, and Sales. It reported 150 milliseconds per request for the Decisions API and 1.6 seconds for the Responses API, a difference of roughly 10.7 times.
The company did not publish enough methodology to reproduce the result, including latency percentiles, payload sizes, concurrency, regions, cache behavior, or network conditions. Developers should treat the figures as vendor benchmarks until independent tests measure the endpoint under production workloads.
OpenAI says the endpoint runs on GPT-6 Luna, the lowest-priced model in the GPT-6 family. It has not documented a separate decision-specific model or explained which serving optimizations produce the latency gain. A constrained answer space can reduce generation work, but attributing the full improvement to constrained decoding would be speculative.
Where the endpoint fits
Preview users applied the API to several recurring tasks:
- Route requests to a model, tool, agent, or business queue.
- Assign labels, rankings, or scores to large collections of records.
- Compare images or identify useful frames in video.
- Select interface controls and form actions from screenshots.
- Flag risky tool calls for deeper review.
- Categorize datasets for subsequent analysis.
Frequent decision calls can dominate an agent’s control loop because each step waits for the previous result. Using OpenAI’s reported figures, 10 serial calls would consume about 1.5 seconds of model time through Decisions and 16 seconds through Responses, before network and application overhead. Actual savings will depend on concurrency, input size, retries, and the number of sequential steps.
Early tests show a close race
Jev, a startup focused on dedicated decision models, provides the clearest direct comparison. Every, which had preview access to OpenAI’s endpoint, reported that Decisions selected the correct control in 76 of 78 steps during an offline replay of computer-use tasks. Jev selected 73 correctly. In a separate thread-classification test, the systems tied on accuracy and Jev returned results faster. Every also reported that Jev performed better across its broader evaluation set.
The small samples and task-specific methods prevent a general quality ranking. They indicate that OpenAI’s endpoint can compete with a purpose-built model on some workloads. OpenAI also offers an operational advantage for existing customers because Decisions uses the same SDK, authentication, and billing infrastructure as its other APIs.
The beta leaves key blanks
| Question | Status at publication | Why developers need it |
|---|---|---|
| Decision-specific pricing | No pricing schedule published | High-volume routing may involve millions of calls |
| Candidate-answer limit | No maximum published | Large taxonomies may require staged classification |
| Fine-tuning | No customization option announced | Specialized domains may need adaptation to internal labels |
| Version controls | Controls were not described in the launch material | Stable versions support reproducible evaluations and audits |
GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens through Responses and Chat Completions. Those rates do not establish the price of Decisions. A per-call fee or latency premium could change the break-even point between the hosted endpoint and a fine-tuned small model running on dedicated infrastructure.
How to test it before production
- Build a representative evaluation set. Include routine cases, ambiguous inputs, rare classes, and costly failure modes.
- Compare against the current system. Measure task accuracy, cost-weighted errors, throughput, and total cost per successful decision.
- Measure tail latency. Record p50, p95, and p99 latency across expected regions, payload sizes, and concurrency levels.
- Validate confidence thresholds. Plot confidence against observed accuracy and route uncertain cases to a fallback model, explicit review option, or human queue.
- Log enough context to debug failures. Capture the question, candidate set, output, confidence, model version when available, latency, and downstream outcome under the application’s privacy rules.
- Design a failure path. Define behavior for timeouts, unavailable endpoints, malformed inputs, retries, and low-confidence results.
A faster control loop for agents
The public beta gives agent builders a dedicated endpoint for frequent control-flow judgments that can be expressed as predicates, finite options, or numeric ranges. Decisions fits routing, gating, classification, and scoring calls. The Responses API remains available for generated prose and open-ended reasoning over long context.
The endpoint’s production value will depend on measured accuracy, tail latency, answer-set limits, and final pricing. Developers can access the beta through the official developer documentation.