OpenAI's Decisions API Gives Developers a Constrained GPT-6 Luna Router
OpenAI's new Decisions API turns GPT-6 Luna into a real-time classifier that picks answers from a fixed list you define.
- OpenAI launched the Decisions API at DevDay 2026, powered by new GPT-6 Luna model.
- Luna picks from a developer-defined finite answer set given text or image context.
- Built for classification, request routing, and choosing an agent's next action.
- Currently in limited preview for selected API customers, with broad release in the coming days.
- Complements the upgraded Agents API with computer use and Ultrafast tier at 300 tokens/sec.
- Signals a shift toward task-specialized GPT-6 variants like Astra, Sol, and Luna.
OpenAI’s Decisions API gives developers a constrained GPT-6 router
OpenAI unveiled the DevDay recap Decisions API at DevDay 2026. The new endpoint uses GPT-6 Luna, a model variant tuned for real-time decisions with a fixed answer set. Developers provide a question, allowed answers, and text or image context; Luna returns one of those answers for the application to consume directly.
The launch gives developers their first access to Luna, a distinct member of the GPT-6 family alongside Astra and Sol. Its narrow interface targets classification, routing, and agent-control tasks that otherwise require prompt constraints, structured-output validation, and recovery logic around a general-purpose model.
One call, one allowed choice
Each Decisions API request defines the decision Luna must make and the answers it may select. The finite answer space lets an application validate the result against the same categories it supplied, without parsing a free-form response.
| Request element | Purpose |
|---|---|
| Question | Defines the decision to make |
| Allowed answers | Sets the choices Luna may return |
| Context | Provides relevant text or images |
| Result | Returns one answer from the declared set |
OpenAI’s example routes a support ticket by sending its contents with a list of eligible teams. Luna evaluates the request and returns the selected team. The underlying task is classification, with a GPT-6-family multimodal model evaluating the context.
Why specialization trims the stack
General-purpose chat models can already classify requests, but production integrations often need prompts that prohibit extra text, schemas that constrain the response, parsers that extract the label, and retries for invalid categories. Luna folds the allowed answer set into the API contract, reducing the amount of application code required to enforce valid output.
A valid label can still be incorrect when an input is ambiguous, adversarial, or outside the supplied taxonomy. Applications that require abstention should include an explicit option such as needs_review or unknown, then route that result to a separate workflow.
Likely workloads
- Routing support tickets, email, and sales leads to queues or teams
- Assigning content to a defined set of moderation labels
- Selecting an agent’s next tool or workflow step from an allowlist
- Classifying product images, forms, receipts, and other documents
- Making frequent branching decisions inside multi-model pipelines
Luna’s place in GPT-6
OpenAI also announced GPT-6.1 Sol, an upgrade focused on agentic coding, computer use, and professional work. The company says Sol delivers performance near Astra at one-fifth of Astra’s standard input and output token prices. An Ultrafast tier offers up to eight times faster generation in Codex, reaching 300 tokens per second, and up to six times faster generation through the API.
Luna serves a different workload within that lineup: bounded selections made repeatedly and quickly. A larger model can plan a task or generate an answer, while Luna chooses the next tool, destination, or policy label from a declared set.
A natural fit for agent loops
The Decisions API can complement the upgraded Agents API guide, which now covers agents that operate software through computer use. Those systems repeatedly choose among actions such as clicking, typing, calling a tool, requesting approval, or stopping. Luna provides a constrained decision point for workflows that expose those actions as allowed choices.

The preview leaves key blanks
The Decisions API is available to selected API customers in limited preview. OpenAI says a broader release will follow in the coming days, but it has not published Luna’s pricing, rate limits, latency measurements, regional availability, or maximum number of choices per request.
Production suitability will depend on latency and classification quality under real traffic. Routing and agent-control systems may call the endpoint at every branch, so tail latency matters alongside average response time. Evaluations should also measure how often Luna confuses neighboring labels, selects a fallback category, or changes its answer when context contains irrelevant or hostile instructions.
Teams testing the preview can shadow existing classifiers, record a confusion matrix for each taxonomy, and compare cost and latency at realistic request volumes. Those results will show whether Luna can replace prompt-and-parse pipelines, conventional classifiers, or fine-tuned models for a given workload.