Liquid AI's d1 Makes Decisions Without Generating a Single Token

Liquid AI's new d1 model returns typed decisions with calibrated probabilities in one call, generating zero output tokens for classification and routing.

·
·
Liquid AI's d1 Makes Decisions Without Generating a Single Token
  • Liquid AI launches d1, its first decision model, taking #1 on Hugging Face's Jev Decision Index
  • Returns typed decisions with calibrated probabilities in one call, generating zero output tokens
  • Three primitives: Noul (yes/no), Choice (pick one), Score (rate on ordered scale)
  • Improvements over prior models: multilingual evals, prompt injection robustness, longer inputs
  • Available now via Liquid API with a free tier; OpenRouter coming soon
  • Cookbook demo: Road Decider game shows real-time frame-by-frame decisioning

Liquid AI’s d1 returns typed decisions without generating tokens

Liquid AI has released d1 documentation for its first decision model, a specialized alternative to language models for classification, routing, approval gates, and scoring inside production software. Liquid reports that d1 leads the Jev Decision Index, which evaluates decision models on a frozen suite of 132,422 requests across 37 benchmarks.

The model targets LLM calls that generate JSON only to express a label, boolean, option, or numeric score. Instead of decoding text and parsing it into a schema, d1 returns typed answers with probabilities while reporting zero generated output tokens. That design can reduce latency, parsing failures, and output-token charges for workloads that fit its limited response types.

Three answer shapes, one request

Every call accepts a state, supplied as plain text or JSON, and one or more questions. Each question uses one of three answer types:

Type Returns Typical use
Noul A probability from 0 to 1 for a yes-or-no question Approval gates, flags, and thresholded booleans
Choice A probability distribution over named options, plus confidence Classification, routing, and action selection
Score A probability-weighted position on an ordered rubric Relevance, severity, quality, and risk scoring

A Score response can fall between rubric levels. For example, a four-level rubric may return 2.9995 instead of rounding to an integer. Applications can preserve that value, map it to a category, or compare it with a threshold.

One request can mix all three types against the same state. A support system could ask for an intent label, urgency score, and bug-report probability in one round trip rather than issuing three model calls.

The call in Python

The API endpoint is https://api.liquid.ai/decisions/v1/systemone. The following TypeSafe SDK example asks d1 whether a customer message is a complaint:

python
import os

from typesafe_sdk import Noul, TypeSafeClient

client = TypeSafeClient(
    api_key=os.environ["LIQUID_API_KEY"],
    base_url="https://api.liquid.ai",
)

result = client.system_one(
    model="d1:free",
    state="I've been waiting three weeks and nobody replies.",
    questions={
        "is_complaint": Noul(
            instructions="Is this a complaint?"
        )
    },
)

print(result.answers["is_complaint"].noul)  # 0.999

The response includes input-token usage and reports usage.output_tokens as 0. The service still returns structured data, but it does not produce that data through autoregressive token generation.

Specialization removes decoding overhead

Language models generate output token by token, even when the application needs only a boolean or label. d1 skips that decoding stage, which gives the service a more predictable latency profile and removes malformed JSON, missing fields, and invalid enum values from the model-output path.

The API also exposes numerical probabilities instead of asking a chat model to write a confidence label. Liquid describes those probabilities as calibrated, although teams should measure calibration on their own traffic before using them for consequential thresholds.

Liquid reports four improvements over its earlier decision models:

  • Higher multilingual evaluation scores
  • Greater resistance to prompt injection in the state field
  • Better handling of longer inputs
  • Faster structured decisions in software workflows

Prompt-injection resistance does not make untrusted input safe by default. Applications should continue to validate inputs, restrict downstream actions, enforce authorization outside the model, and test attacks that reflect their deployment environment.

The benchmark lead needs local testing

The Jev Decision Index runs 37 benchmarks, with 19 scored benchmarks contributing to its headline composite. Those scored tests are divided equally among five areas: Tools & Automation, Retrieval & Classification, Language Understanding, Knowledge & Reasoning, and Arts & Human Judgment. The final index is 100 times the mean of the five area scores, producing a value from 0 to 100.

Jev 1.13 previously led the index with a composite near 74.4 before d1 took the top position, according to the published leaderboard. The frozen request set makes model comparisons more consistent, but a leaderboard cannot predict production accuracy, calibration, latency, or cost for a specific label set and traffic distribution.

Workloads that match the API

d1 fits workflows whose final output already has a fixed schema:

  • Ticket, message, and email routing
  • Content moderation and safety flags
  • Agent tool-call approval gates
  • Narrow rubric-based evaluations
  • Model routing and inference cascades
  • Fraud, risk, severity, and priority scoring
  • RAG reranking with relevance probabilities

Generative workloads remain outside the model’s scope. Liquid recommends a language model for free-form writing, summarization, code generation, multi-turn conversation, open-ended questions, and complex multi-step reasoning. Any workflow that must produce a new sentence needs a generative model or a separate templating layer.

Production thresholds require evidence

Teams evaluating d1 should test the full decision pipeline rather than compare model scores alone:

  • Define labels precisely: Write instructions and rubrics that distinguish adjacent classes with concrete criteria.
  • Measure on representative data: Include rare classes, multilingual inputs, long states, and adversarial examples.
  • Check calibration: Compare predicted probabilities with observed outcomes for each important segment.
  • Set thresholds by cost: Choose cutoffs based on the consequences of false positives, false negatives, and abstentions.
  • Keep a fallback path: Route uncertain or high-risk cases to rules, a larger model, or human review.
  • Monitor drift: Track class frequency, probability distributions, error rates, latency, and input-token usage after launch.

A game loop demonstrates the response shape

Liquid’s Road Decider example uses d1 as the controller for a pixel-art driving game. For each frame, the model selects LEFT, CENTER, or RIGHT and returns a confidence score.

Split-screen Road Decider game with directional choices and confidence scores
The demo maps a Choice response directly to one of three driving actions.

The example illustrates the API’s intended control pattern: provide the current state, request a fixed set of actions, and consume the resulting probability distribution. Similar patterns appear in agent action selection, robotics gates, moderation pipelines, and model routers, although real-time control systems also need deterministic safety constraints outside the model.

Access, pricing, and rollout

d1 is available through the API console under the d1:free model name. Developers need a Liquid API key, and production users should confirm current rate limits, data-handling terms, regional availability, and paid pricing in the console before deployment.

Liquid says OpenRouter support is planned but has not provided a launch date. Because d1 reports no output tokens, usage is driven primarily by the input state and question definitions; actual cost and latency will still depend on request size, service limits, and workload volume.

d1 gives developers a narrower interface for decisions that many applications currently obtain through generated JSON. Its benchmark position supports further evaluation, while a production trial should determine whether the model’s accuracy, calibration, latency, and operating cost hold on the application’s own data.

Trending
  • No trending articles

Comments

avatar

Next Reads