Inception Releases Mercury Decide to Make AI Routing 14x Faster
Inception's new Mercury Decide is a structured decision model that returns typed answers with calibrated probabilities at 14 decisions per second, free on OpenRouter.
- Inception launched Mercury Decide, a structured decision model returning typed answers with calibrated probabilities.
- Runs up to 14 decisions per second on a dedicated
/v1/systemoneendpoint, not chat-compatible. - Returns a choice, score, or yes/no answer with probability read from the model directly.
- Output tokens are free because probabilities are extracted from internal distributions, not generated as text.
- Free for early access on OpenRouter with a 33K context window and single-provider hosting.
- Built for triage, routing, moderation, agent step selection, and evaluation loops rather than text generation.
Mercury Decide targets fast, bounded decisions
Inception has released Mercury Decide, a hosted model for structured decisions. A request supplies contextual state, evaluation criteria, and an expected answer type. The model returns a choice, numeric score, or Boolean answer with a confidence probability.
Inception calls the service a System One endpoint, borrowing Daniel Kahneman’s term for fast judgment. The design targets moderation, routing, grading, fraud triage, and other workflows that need a bounded answer without generated prose.
One request, three answer shapes
Each request combines three concepts:
- State: The facts or context the model should evaluate.
- Question: The decision the model must make.
- Rubric: The permitted answer format and evaluation criteria.
The dedicated /v1/systemone endpoint supports three response types:
| Type | Returned value | Typical use |
|---|---|---|
| Choice | One option from an enumerated set | Tool selection or request routing |
| Score | A numeric value | Grading, ranking, or risk assessment |
| Boolean | Yes or no | Policy checks or escalation gates |
Structured fields remove the need to prompt for JSON and parse generated text. Inception says the accompanying probability comes directly from the model’s predictive distribution. The endpoint therefore incurs no output-token charge during early access.
Fast enough for the routing layer
OpenRouter advertises throughput of up to 14 decisions per second. Actual latency will depend on request size, network conditions, provider load, and concurrency, so the published figure does not guarantee end-to-end performance in an application.
Inception attributes the speed to diffusion language modeling, the same broad approach used by Mercury 2.5. Diffusion models refine multiple positions in parallel across several passes. Autoregressive language models generate tokens sequentially. A task that emits one bounded value gives the diffusion system little output to construct.
Strong fits have hard boundaries
Mercury Decide suits workflows with a fixed answer space and enough context to state the decision clearly. Likely applications include:
- Moderation and safety-policy classification
- Fraud or anomaly triage for transaction streams
- Agent tool selection from a fixed menu
- Evaluation of model outputs against a rubric
- Routing requests to specialized models or human reviewers
Its output is limited to choices, scores, and Boolean values. Writing, summarization, code generation, and open-ended analysis require a generative model. Tasks with unstable labels or poorly defined criteria also weaken the value of a bounded decision endpoint.
OpenRouter describes Mercury Decide as leading JevBench, which evaluates models that map supplied state and a bounded rubric to a typed answer. That benchmark matches the product’s intended workload, but it cannot establish accuracy, latency, or calibration for a specific production dataset.
Calibration carries the load
Internal probabilities are useful only when they remain calibrated. Among decisions assigned 80 percent confidence, roughly 80 percent should be correct over a representative sample. Reading a probability from the model’s distribution avoids a generated self-rating, but that mechanism alone does not prove calibration.
A production evaluation should measure accuracy and calibration separately for each decision type, threshold, language, and data segment. It should also track drift after prompt, rubric, model, or traffic changes. A practical deployment can then:
- Send every eligible event to Mercury Decide.
- Accept high-confidence answers that meet a tested threshold.
- Route uncertain cases to a larger model or human reviewer.
- Audit sampled decisions and update thresholds as performance changes.
Early access requires custom integration
| Availability | Early access through OpenRouter |
|---|---|
| Price | Free during early access, with no output-token charge |
| Context window | 33K tokens |
| Provider | One provider, hosted directly by Inception |
| Endpoint | /v1/systemone |
| Compatibility | No OpenAI-compatible chat-completions path |
Existing chat SDK calls will not map directly to the Decisions API. Teams will need request code for the System One endpoint, response validation for each answer type, timeout and retry handling, and monitoring for confidence distributions. Early-access pricing, limits, and provider availability may change.
A smaller primitive for production AI
Many production systems use general-purpose language models for narrow judgments, then pay for generated explanations, JSON parsing, and retries when formatting fails. Mercury Decide packages the bounded judgment as a dedicated call with typed output and an explicit probability.
The model could serve as a hosted alternative to a custom classifier when labels vary by workflow or training data is scarce. Its practical value will depend on domain accuracy, stable calibration, latency under concurrent load, and the cost that follows early access. Those measurements will determine whether it can reliably handle the first pass while expensive models and human reviewers focus on ambiguous cases.