Liquid AI's d1-3B Makes Structured AI Decisions in 8 Milliseconds Without Generating Text
Liquid AI's d1-3B skips token generation entirely, returning calibrated classifications and scores in a single forward pass, with multimodal input on edge hardware.
- Liquid AI released d1-3B, an open-weight 3.1B multimodal decision model that answers in one forward pass with zero output tokens.
- Scores 48.57 on Decision Index 0.2.1, beating every sub-10B model and matching Decider 35B-A3B.
- Latency: 8 ms on RTX 4090, 9 ms on AMD MI325X, 16 ms on Jetson AGX Thor, 50 ms on Jetson Orin Nano.
- Supports three primitives: Noul (yes/no with probability), Choice (named options), Score (ordered rubric).
- Built on LFM2.5-VL-3B via weight averaging, multi-seed fine-tuning, and checkpoint merging.
- Available on Hugging Face, with vLLM, SGLang, and blog post covering deployment.
Liquid AI’s d1-3B makes decisions in one pass
Liquid AI has released d1-3B, an open-weight multimodal model that returns structured predictions in a single forward pass and generates zero output tokens. Built on Liquid Foundation Models, it removes the sequential decoding loop used by generative transformers. Pipelines that need a yes-or-no answer, category, ranking, or score can therefore avoid generating and parsing text.
The 3.1B-parameter model builds on LFM2.5-VL-3B. Liquid also released an experimental 600M sibling called d1-omni-600M. Both models are available on Hugging Face.
One pass, three answer types
Each request contains a state, such as text, JSON, an image, or a combination, plus a dictionary of named questions. The model reads the state once and returns typed answers with probabilities that Liquid describes as calibrated. The API defines three primitives:
- Noul: Answers a yes-or-no question with a probability from 0 to 1. For example, “Is this message spam?” might return 0.92.
- Choice: Selects one named option and returns a probability distribution across all supplied options.
- Score: Rates the input against an ordered rubric and returns a probability-weighted position on that scale.
A single API call can apply several decisions to the same customer message:
questions = {
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?",
},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorized use",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": [
"Can wait",
"Today",
"Blocking the customer now",
],
},
}
model.system_one(
"I was charged twice this month, please refund one of them.",
questions,
)This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.