LiquidAI's d1 Hits Vercel With Zero Output Token Costs
Liquid AI's zero-output-token decision model lands on Vercel's AI Gateway, giving developers a cheap drop-in for classification, routing, and scoring work.
- Liquid AI's d1 decision model is now available on Vercel's AI Gateway.
- Pricing is $0.04 per million input tokens and $0 for output, since d1 generates no tokens.
- Returns calibrated probabilities over typed Choice, Score, and yes/no (Noul) questions in one call.
- 66K context window, aimed at classification, routing, moderation, and LLM-as-judge workloads.
- Claims first-place finish over Jev on Hugging Face's Decision Index, with better prompt injection resistance.
- API-only today; Liquid has hinted at downloadable weights but no release date yet.
Liquid AI’s d1 decision model is now available through Vercel’s AI Gateway, extending access beyond Liquid’s API console and OpenRouter. Each call evaluates structured input and returns probabilities over answers defined by the developer. The model generates no output tokens, which changes the cost profile for classification, routing, moderation, and scoring workloads.
| Specification | Vercel listing |
|---|---|
| Input price | $0.04 per million tokens |
| Output price | $0 per million tokens |
| Context window | 66K tokens |
| AI SDK interface | experimental_evaluate |
| Deployment | Hosted API only |
A scorer with three answer types
A decision model handles questions whose possible answers are known before the request. Calibration means a probability of 0.9 should correspond to roughly nine correct decisions out of ten across comparable cases. That property lets applications set explicit thresholds for automated actions, manual review, or fallback processing.
Liquid exposes three question types that can be combined in one request:
- Noul: Answers a yes-or-no question with a probability between 0 and 1. Liquid’s example, “Is this message a complaint?”, returned 0.999.
- Choice: Selects from a named set and returns the leading option, the full probability distribution, and a confidence value.
- Score: Evaluates input against an ordered rubric and returns a probability-weighted position. The levels use zero-based indexes, so a four-level urgency rubric spans 0 through 3.
Vercel routes d1 calls through the AI SDK’s experimental evaluate function:
import { experimental_evaluate as evaluate } from 'ai';
const result = await evaluate({
model: 'liquid/d1',
state: 'The support agent issued a full refund to the customer.',
questions: {
refunded: {
type: 'boolean',
instructions: 'Was a refund issued?',
},
},
});The request supplies the state to evaluate and a typed set of questions. The response contains numeric decisions, while usage.output_tokens remains 0. Applications that require summaries, explanations, or other original text still need a generative model.
Why d1 is showing up now
Liquid added d1 to its own API several days before the Vercel listing and introduced it as the company’s first decision model. Liquid also says d1 is the first model to outperform Jev on Hugging Face’s Decision Index, a benchmark for fixed-answer decision tasks. Jev has served as a reference model for this category and has prompted integrations with LLM applications as well as development of similar specialized models.
Liquid reports gains on multilingual evaluations, long-input handling, and resistance to prompt injection. Injection resistance matters when a model evaluates untrusted content, such as support tickets, user posts, retrieved documents, or proposed agent tool calls. An attacker can place instructions inside that content to influence the decision, so teams should test d1 against application-specific attacks before relying on benchmark results.
The output-token line drops to zero
General-purpose models often generate labels, JSON wrappers, explanations, or reasoning tokens around a classification result. d1 derives a probability distribution from its internal scores and returns the numeric answer directly, eliminating generated-output charges. At Vercel’s listed rate, processing 10 million input tokens would cost $0.40 in model input fees, excluding any additional platform charges.
The savings can accumulate in agent systems that run several checks for each user request. A single turn might require intent routing, safety classification, tool approval, and response grading. Each d1 call incurs input-token costs while adding no output-token bill.
Probability distributions also support clearer control logic than free-text confidence claims. A team could approve tool calls above 0.95, send results between 0.70 and 0.95 for review, and reject lower scores. Those thresholds should be tuned against representative production data because calibration can vary by task, language, and input distribution.
API access brings trade-offs
d1 is currently API-only and cannot be fine-tuned through Liquid’s model library. No GGUF, MLX, or ONNX weights are available for local deployment. Liquid has indicated that downloadable weights are planned, but it has provided no release date. Organizations that require on-premises inference or prohibit external processing cannot deploy the model under the current terms.
Vercel’s integration also uses the AI SDK’s experimental_evaluate interface, which may change. Production users should pin the SDK version, isolate the call behind an internal adapter, and test response parsing during upgrades.
Teams already sending classification or routing requests to generative models can compare d1 through shadow traffic. A useful evaluation should measure task accuracy, probability calibration, latency, multilingual performance, injection resistance, and total request cost. Workloads with fixed answer sets are the clearest candidates; workloads that compose new text remain outside d1’s scope.