Inception Releases Mercury Decide to Make AI Routing 14x Faster

Inception's new Mercury Decide is a structured decision model that returns typed answers with calibrated probabilities at 14 decisions per second, free on OpenRouter.

·
·
Read4 min
TypeNews
TopicLlms · Api
  • Inception launched Mercury Decide, a structured decision model returning typed answers with calibrated probabilities.
  • Runs up to 14 decisions per second on a dedicated /v1/systemone endpoint, not chat-compatible.
  • Returns a choice, score, or yes/no answer with probability read from the model directly.
  • Output tokens are free because probabilities are extracted from internal distributions, not generated as text.
  • Free for early access on OpenRouter with a 33K context window and single-provider hosting.
  • Built for triage, routing, moderation, agent step selection, and evaluation loops rather than text generation.

Mercury Decide targets fast, bounded decisions

Inception has released Mercury Decide, a hosted model for structured decisions. A request supplies contextual state, evaluation criteria, and an expected answer type. The model returns a choice, numeric score, or Boolean answer with a confidence probability.

Inception calls the service a System One endpoint, borrowing Daniel Kahneman’s term for fast judgment. The design targets moderation, routing, grading, fraud triage, and other workflows that need a bounded answer without generated prose.

One request, three answer shapes

Each request combines three concepts:

  • State: The facts or context the model should evaluate.
  • Question: The decision the model must make.
  • Rubric: The permitted answer format and evaluation criteria.

The dedicated /v1/systemone endpoint supports three response types:

Type Returned value Typical use
Choice One option from an enumerated set Tool selection or request routing
Score A numeric value Grading, ranking, or risk assessment
Boolean Yes or no Policy checks or escalation gates

Structured fields remove the need to prompt for JSON and parse generated text. Inception says the accompanying probability comes directly from the model’s predictive distribution. The endpoint therefore incurs no output-token charge during early access.

Fast enough for the routing layer

OpenRouter advertises throughput of up to 14 decisions per second. Actual latency will depend on request size, network conditions, provider load, and concurrency, so the published figure does not guarantee end-to-end performance in an application.

Inception attributes the speed to diffusion language modeling, the same broad approach used by Mercury 2.5. Diffusion models refine multiple positions in parallel across several passes. Autoregressive language models generate tokens sequentially. A task that emits one bounded value gives the diffusion system little output to construct.

Strong fits have hard boundaries

Mercury Decide suits workflows with a fixed answer space and enough context to state the decision clearly. Likely applications include:

  • Moderation and safety-policy classification
  • Fraud or anomaly triage for transaction streams
  • Agent tool selection from a fixed menu
  • Evaluation of model outputs against a rubric
  • Routing requests to specialized models or human reviewers

Its output is limited to choices, scores, and Boolean values. Writing, summarization, code generation, and open-ended analysis require a generative model. Tasks with unstable labels or poorly defined criteria also weaken the value of a bounded decision endpoint.

OpenRouter describes Mercury Decide as leading JevBench, which evaluates models that map supplied state and a bounded rubric to a typed answer. That benchmark matches the product’s intended workload, but it cannot establish accuracy, latency, or calibration for a specific production dataset.

Calibration carries the load

Internal probabilities are useful only when they remain calibrated. Among decisions assigned 80 percent confidence, roughly 80 percent should be correct over a representative sample. Reading a probability from the model’s distribution avoids a generated self-rating, but that mechanism alone does not prove calibration.

A production evaluation should measure accuracy and calibration separately for each decision type, threshold, language, and data segment. It should also track drift after prompt, rubric, model, or traffic changes. A practical deployment can then:

  1. Send every eligible event to Mercury Decide.
  2. Accept high-confidence answers that meet a tested threshold.
  3. Route uncertain cases to a larger model or human reviewer.
  4. Audit sampled decisions and update thresholds as performance changes.

Early access requires custom integration

Availability Early access through OpenRouter
Price Free during early access, with no output-token charge
Context window 33K tokens
Provider One provider, hosted directly by Inception
Endpoint /v1/systemone
Compatibility No OpenAI-compatible chat-completions path

Existing chat SDK calls will not map directly to the Decisions API. Teams will need request code for the System One endpoint, response validation for each answer type, timeout and retry handling, and monitoring for confidence distributions. Early-access pricing, limits, and provider availability may change.

A smaller primitive for production AI

Many production systems use general-purpose language models for narrow judgments, then pay for generated explanations, JSON parsing, and retries when formatting fails. Mercury Decide packages the bounded judgment as a dedicated call with typed output and an explicit probability.

The model could serve as a hosted alternative to a custom classifier when labels vary by workflow or training data is scarce. Its practical value will depend on domain accuracy, stable calibration, latency under concurrent load, and the cost that follows early access. Those measurements will determine whether it can reliably handle the first pass while expensive models and human reviewers focus on ambiguous cases.

Trending
  • No trending articles

Comments

avatar

Next Reads