AutoTrust AI Turns Gemma 4 Into a 45ms Calibrated Decision API
AutoTrust AI Lab released GEV-26B-Decide, a Gemma-4 based decision model that outputs calibrated probabilities and thinks only when uncertain.
- AutoTrust AI Lab released GEV-26B-Decide, a Gemma-4-26B decision model with calibrated option probabilities.
- System 1 answers in ~45 ms per decision; adaptive mode invokes reasoning only below 0.8 confidence.
- Big gains: GPQA Diamond 42.9 to 78.6, BBH 75.0 to 92.0, MMLU-Pro 65.0 to 84.6.
- Browser agent demo completes 95% of multi-step tasks at ~85 ms per click on one B200.
- Supports yes/no, 0-5 score, and up to 256-way choice; 256K token context, image inputs.
- Ships with vLLM server,
POST /v1/decideendpoint, Apache-2.0 adapter and decision head.
GEV-26B-Decide turns Gemma 4 into a calibrated decision API
AutoTrust AI Lab has released GEV-26B-Decide, an open-weights model built on Gemma-4-26B-A4B-it. It maps typed questions to probability distributions over fixed options, then invokes step-by-step reasoning when the leading probability falls below a confidence threshold.
Applications can use those probabilities to route uncertain cases, select actions, and avoid parsing generated labels such as “A” or “B.” Calibration means that predictions assigned 80% confidence should be correct about 80% of the time.
- Decision endpoint: Custom
POST /v1/decideroute alongside OpenAI-compatible routes - Inputs: Text and images, with contexts up to 256K tokens
- Fast path: One forward pass, reported at about 45 ms
- License: Apache-2.0 for the adapter, decision head, and calibration files
- Base model: Covered separately by the Gemma 4 terms
One backbone, two inference paths
AutoTrust labels the two paths System 1 and System 2. System 1 reads a probability distribution from the model’s hidden state in one forward pass. System 2 runs the same loaded backbone in Gemma 4’s thinking mode, using the original state, question, and options as context. Both paths support text and image inputs through one vLLM engine.
System 1 accepts three decision schemas:
| Kind | Purpose | Output space |
|---|---|---|
noul |
Binary questions | Yes or no |
choice |
Option selection | 2 to 256 choices |
score |
Rating | 0 to 5 |
Each request returns a probability for every available option. System 1 performs no token generation, which removes decoding latency and avoids extracting an answer from free-form text.
Uncertainty buys the extra compute
Adaptive inference uses the fast path unless its confidence gate triggers:
- Read the first distribution. System 1 produces
p1in one forward pass. - Apply the confidence gate. System 2 runs when the highest probability in
p1falls below the default threshold of0.8. - Reason within a budget. The
think_budgetsetting caps the number of thinking tokens. - Read and combine. When the thinking channel closes, the model reads a second distribution,
p2, and returnsp = 0.5p1 + 0.5p2.
The 50/50 blend preserves information from the calibrated fast path because
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.