System One models are carving out a new layer in the AI stack
TypeSafe's Jev returns a typed answer in 70 to 500 ms at $0.042 per million input tokens, with no metered output-token cost
- TypeSafe reports Jev at 70-500 ms end to end and $0.042 per million input tokens, with no metered output-token cost.
- TypeSafe's table puts System One shaped queries at 40-200x faster than frontier LLMs at similar intelligence, and the 444.6x cheaper figure is the workflow-eval high end.
- Laya has 421 million parameters, and its developers report about 33 ms per question on the multilingual checkpoint on a Tesla T4.
- In the CLM team's zero-shot tests, CLM ran up to 9x faster than Jev.
- Fine-tuned CLM scored 81.6% on 38 held-out DeepSWE tasks and 87.6% on 30 Terminal-Bench 2.1 tasks, versus 71.1% and 83.1% for Jev.
LLMs have become the default building block for almost every AI task, from writing code to deciding which support queue should receive an email.
But many of those tasks don't require generation at all. They require fast, reliable decisions, and using a full generative model for every one can become expensive as applications scale.
Jev and a growing group of "System One" models are taking a different approach.
Today we break down where they fit in your AI stack, and when developers should use them instead of code or an LLM.
Where Jev and System One models fit in your AI stack
TypeSafe AI released Jev, its new "System One Model," on September 15. Jev is designed to make lightning-fast structured decisions instead of generating text like LLMs.
The launch drew strong developer interest, and open alternatives followed quickly. Laya provides an open System One-style model, while Stanford and Nvidia researchers released CLM-8B, which tackles the same class of bounded decisions with a different contrastive architecture.
It is too early to call System One an established model category, but the rapid appearance of alternatives suggests Jev has exposed a useful niche: models built to choose, score and verify instead of generate.
AI applications often make decisions that are too fuzzy for hand-written rules but do not require a generative model: choosing an agent tool, routing a request, enforcing a guardrail, or selecting the best answer from several candidates.
System One models provide another building block for handling those decisions, with the potential to make production AI systems more predictable and economical as they scale.
How Jev changes the model interface
Large language models are autoregressive. They create output one token at a time, which is good for generating text and code. But that flexibility is unnecessary if you only want to decide whether a support ticket belongs to billing, sales or technical support.
Jev handles these kinds of situations. You provide it with the application state, questions and possible answers. It returns typed values with probability and confidence information. Instead of asking an LLM to generate and then parse "This looks like a billing issue," your application can directly receive a typed "billing" value.
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves