Liquid AI Opens d1 Decision Models That Classify Without Generating a Single Token
Liquid AI drops two open-weight decision models that classify, route, and score text, images, and audio in a single forward pass with no generated tokens.
- Liquid AI released Open d1: d1-3B (text+vision) and d1-omni-600M (text+image or text+audio), both open-weight.
- Decision models return calibrated probabilities over fixed outcomes in a single forward pass, with zero generated tokens.
- d1-3B hits 8 ms on an RTX 4090, 16 ms on Jetson AGX Thor, 50 ms on Jetson Orin Nano.
- Scores 48.57 on Decision Index v0.2.1, matching a decision model 12x its size.
- Three question primitives: Noul (yes/no), Choice (pick one), Score (ordered rubric), composable in one call.
- Day-one llama.cpp support across Apple, AMD, Qualcomm, NVIDIA; weights on Hugging Face.
Liquid AI open-sources two low-latency decision models
Liquid AI has open-sourced two members of its d1 family: d1-3B, which accepts text and images, and the experimental d1-omni-600M, which accepts text paired with either images or audio. These decision models score predefined outcomes in one forward pass, returning calibrated probabilities while generating zero tokens. The design targets classification, routing, scoring, and guardrail checks that need predictable output at millisecond-scale latency.
One pass, three question types
A conventional LLM classifier generates an answer token by token, often followed by JSON parsing, schema validation, and retries. A d1 model computes probabilities for outcomes supplied with the request and returns structured values directly.
Each request includes an input state and one or more questions built from three primitives:
- Noul: A yes-or-no question that returns a probability between 0 and 1. It suits gates such as spam detection or tool-call approval.
- Choice: A selection from named options that returns a probability distribution across them. It can route a support ticket to billing, technical support, or account management.
- Score: A rating against an ordered rubric that returns a probability-weighted position on the scale. It can estimate severity, urgency, or response quality.
Multiple questions can share the same state and inference request. A support pipeline could classify intent, select a team, and score urgency together, reducing the overhead of separate model calls.
Latency across five accelerators
Liquid reports the following single-question latency for d1-3B. The MI325X figure also includes throughput for packed states, a batching measure reported by the company.
| Hardware | Latency | Additional result |
|---|---|---|
| NVIDIA GeForce RTX 4090 | 8 ms | |
| AMD MI325X | 9 ms | 1,106 packed states per second |
| NVIDIA Jetson AGX Thor | 16 ms | Three questions in 20 ms |
| NVIDIA Jetson AGX Orin | 26 ms | |
| NVIDIA Jetson Orin Nano | 50 ms |
The AGX Thor result shows the benefit of sharing computation across questions: tripling the question count increased reported latency from 16 ms to 20 ms. That scaling can support several routing or guardrail checks against each video frame or application state.
Evidence is strongest on text
On the public split of Decision Index v0.2.1, a composite benchmark for decision models, d1-3B scores 48.57. Liquid reports that it leads models below 10 billion parameters and performs comparably to Decider 35B-A3B, which the company describes as 12 times larger.
Across seven public text benchmarks covering reading comprehension, toxicity detection, intent classification, medical question answering, and cross-lingual understanding, d1-3B posts a mean score of 82.9. Decider 4B scores 81.1 on the same collection.
Liquid describes d1-omni-600M as an experimental checkpoint. It scores 15.95 on Decision Index v0.2.1, along with 95.8 on Civil Comments toxicity detection and 79.5 on PAWS-X paraphrase identification. Its 600 million parameters make it the smaller option for memory-constrained moderation and routing workloads.
The published evaluation leaves gaps around its newer modalities. Liquid did not report results for the private vision split, and dedicated audio decision benchmarks remain undeveloped. Evidence for image and audio quality therefore comes mainly from demonstrations rather than broad benchmark coverage.
Two backbones and a hands-on recipe
d1-3B and d1-omni-600M use different architectures suited to their parameter counts and supported inputs.
| Model | Base | Architecture | Inputs |
|---|---|---|---|
| d1-3B | LFM2.5-VL-3B | Decoder-only vision-language model | Text and images |
| d1-omni-600M | LFM2.5-Encoder-350M | Bidirectional encoder with vision and audio components | Text with images or audio |
Liquid’s training notes describe model merging, varied fine-tuning runs, and targeted multimodal adapters:
- For d1-3B, the team averaged LFM2.5-2.6B with the text backbone of LFM2.5-VL-3B, fine-tuned several checkpoints with different seeds and data mixtures, and merged the resulting models.
- Long training inputs, shuffled answer options, and removal of shortcuts in the data produced larger gains than the more complex methods Liquid tested.
- For d1-omni-600M, the team connected a FastConformer audio encoder through an adapter, then added the vision encoder from LFM2.5-VL-450M.
- The vision integration used LoRA, a parameter-efficient fine-tuning method that updates small auxiliary weight matrices. Those updates activated only for image inputs.
One request in Python
Liquid exposes the hosted model through a system_one endpoint. Each call contains a state, typed questions, and optional images. The following example assumes an initialized client and an imported Choice helper:
result = client.system_one(
model="d1",
state="Charged twice for my subscription last month...",
questions={
"intent": Choice(
instructions="Which team should handle this?",
criteria={
"billing": "Charges, refunds, payments",
"technical": "Bugs or how-to questions",
"account": "Login, profile, subscription",
},
),
},
)
print(result.answers["intent"].choice) # "billing"
print(result.answers["intent"].confidence) # 0.9996The response contains the probability distribution, selected answer, and confidence value. Applications can apply explicit thresholds, route high-confidence cases automatically, and send ambiguous inputs to another model or human review.
Workloads that match the output
The fixed-output interface suits production paths currently handled by prompted language models or task-specific classifiers:
- Agent guardrails and tool-call approval, where each action requires a policy check.
- Content moderation and personally identifiable information detection before input reaches a larger model.
- Reranking and model-routing cascades with tight latency budgets.
- Visual inspection against predefined criteria, using Noul questions on camera frames.
- Support-ticket classification, assignment, and urgency scoring in one request.
- Evaluation pipelines that need a stable score schema instead of generated prose.
d1 only scores outcomes defined by the caller and does not produce open-ended text. Free-form generation, multi-turn conversation, summarization, code generation, and open-ended Q&A remain jobs for generative models. Teams using probability thresholds should also validate calibration on their own data and monitor it as traffic patterns and class frequencies change.
From Hugging Face to edge devices
Liquid publishes both checkpoints on Hugging Face under open-weight licenses. The release also includes llama.cpp support for Apple, AMD, Qualcomm, and NVIDIA hardware, including NVIDIA’s NVFP4 format.
The hosted API offers a paid d1 tier with vision support and a text-only d1:free tier. Requests can contain up to eight images and 10,000 patches of 32 by 32 pixels. Liquid bills image processing at 1.5 tokens per patch.
Jetson AI Lab has published model cards for edge deployment. Liquid also provides a live demo space with ten camera-based examples, including gesture-controlled games and content moderation running one forward pass per frame.