Microsoft's Decision-1 Scores AI Choices 35x Faster Than GPT-6

Microsoft's new decision-scoring model handles classification, routing, and verification tasks at 35x the speed of GPT-6 Sol with calibrated probabilities.

·
·
·
Read5 min
TypeNews
  • Microsoft launched Microsoft-Decision-1, a decision-scoring model for classification, routing, and verification tasks.
  • Reported 35x faster than GPT-6 Sol and 4.5x faster than Quyet-1.0-Large at P50 latency.
  • Pricing is $0.042 per million input tokens with free output tokens on Microsoft Foundry.
  • Post-trained from Qwen3.5-9B for single-pass scoring with calibrated per-option probabilities.
  • Flips decisions on only 1.3% of input perturbations across paraphrases, reorderings, and formatting noise.
  • Internal XBOX deployment ran 14x faster and 200x cheaper than GPT-6 Sol on feedback labeling.

Microsoft releases Decision-1 for fast, fixed-choice scoring

Microsoft has released Microsoft-Decision-1, a model that scores options from a predefined list. It targets frequent software decisions such as routing a request, classifying a ticket, grading an agent action, filtering content, or accepting an output.

The model is available through the Foundry catalog, with OpenRouter support planned. Microsoft charges $0.042 per million input tokens and nothing for output tokens.

Detail Microsoft-Decision-1
Base model Post-trained Qwen3.5-9B
Input Task content, fixed answer options, and optional grading criteria
Output A probability score for each option
Execution One forward pass
Input price $0.042 per million tokens
Output price $0 per million tokens

Fixed choices, usable probabilities

Each request supplies a known set of options, and the model returns a probability score for every choice through a structured API response. Supported patterns include yes-or-no decisions, multiple-choice questions, rating scales, rubric-based grading, and evaluations of AI responses or agent actions.

Applications can apply thresholds directly to those scores. A router might process a request automatically above 0.95 confidence, send an uncertain case to a larger model between 0.60 and 0.95, and request human review below 0.60.

One pass trims every branch

Microsoft post-trained Qwen3.5-9B for single-pass decision scoring. Its execution path excludes free-form generation, tool loops, and multi-step reasoning traces, reducing the work required for each classification.

Diagram of the Microsoft-Decision-1 decision-scoring architecture
Microsoft’s diagram of the Decision-1 scoring workflow.

A workflow that makes 20 sequential decisions accumulates two seconds of delay when each scorer contributes 100 milliseconds. Agent pipelines often contain dozens of dependent routing, verification, and grading calls, so per-call latency can shape the response time of the entire system.

Fastest in Microsoft’s test

Microsoft evaluated Decision-1 on 36 benchmarks containing nearly 150,000 questions that the company says were excluded from training. The model recorded the highest accuracy and lowest latency among the systems in that comparison.

Measure Reported result
Benchmark scope 36 benchmarks and nearly 150,000 questions
Accuracy Highest among the compared models
Speed against Quyet-1.0-Large 4.5 times faster
Speed against GPT-6 Sol 35 times faster
Average decision-flip rate 1.3% across eight input perturbations

Microsoft also tested sensitivity to irrelevant changes in the input. Decision-1 produced no decision flips when researchers paraphrased option descriptions, reversed the options, or shuffled their order, addressing position and wording effects that can make model-based classifiers unstable.

These figures are vendor-reported results, and production performance will depend on each application’s prompts, label distribution, hardware, network path, and decision thresholds. A representative evaluation set remains necessary before replacing an existing classifier or router.

Confidence becomes a control signal

Microsoft describes the returned probabilities as calibrated scores. A 0.9 score is well calibrated when predictions receiving that score are correct about nine times out of 10 across representative cases.

Reliable confidence estimates let software act automatically on clear cases and escalate uncertain ones. Teams evaluating the model should measure calibration on their own data with reliability plots, expected calibration error, or Brier scores, then set thresholds according to the cost of false positives and false negatives.

Four deployments inside Microsoft

Team Workload Reported result
Xbox Research Sorted more than 10,000 survey responses and reviews from Steam and X into predefined themes Quality competitive with GPT-6 Sol, more than 14 times faster, and 200 times less expensive
Copilot Graded the quality of chat and agent responses Quality competitive with GPT5.6 Luna and 100 times faster
On-call engineering Selected relevant knowledge from logs, tickets, calls, messages, and other incident data Better retrieval quality and speed than the unnamed LLM baseline
Microsoft Discovery Graded experiments for an agent that revises and repeats its plan 46 times more consistent scoring at three times the speed, with adaptive replanning nearly four times faster

All four examples come from Microsoft, and several comparisons omit full baseline and evaluation details. They show the workloads the company is targeting, while independent tests will be needed to establish how broadly the reported gains transfer.

Good fits and hard limits

  • Agent verification: continue, retry, stop, or hand off.
  • Model routing: select a model from the request type or risk level.
  • Intent classification: map a request to a known category.
  • Evaluation: grade responses against a rubric.
  • Search and retrieval: score candidate results for relevance.
  • Safety filtering: classify content against defined policies.
  • Computer-use agents: choose the next action from available controls.
  • Data labeling: assign records to an established taxonomy.

Tasks requiring free-form writing, open-ended planning, tool execution, or an explanatory reasoning trace fall outside the model’s scope. Every request must reduce to predefined options or rubric-based scores.

Benchmark it on your own traffic

  1. Build an evaluation set that reflects production inputs, rare cases, and the real label distribution.
  2. Compare accuracy, calibration, latency, and cost with the current router or classifier.
  3. Shuffle and paraphrase options to test sensitivity to wording and order.
  4. Set separate confidence thresholds for automatic action, escalation, and rejection.
  5. Measure end-to-end latency, including network and orchestration overhead.
  6. Monitor calibration and class performance as traffic changes.

A dedicated layer for agent decisions

Microsoft is positioning decision models as a distinct category for structured outputs that software can consume immediately. In an agent stack, a generative model can handle open-ended work while Decision-1 handles repeated routing, grading, filtering, and verification calls.

Microsoft plans to release versions based on its MAI models and OpenAI models. Developers can keep the fixed-choice integration pattern while evaluating whether later backbones improve accuracy, calibration, or coverage for their workloads.

Trending
  • No trending articles

Comments

avatar

Next Reads