AutoTrust AI Turns Gemma 4 Into a 45ms Calibrated Decision API

AutoTrust AI Lab released GEV-26B-Decide, a Gemma-4 based decision model that outputs calibrated probabilities and thinks only when uncertain.

·
·
·
AutoTrust AI Turns Gemma 4 Into a 45ms Calibrated Decision APIPRO
Read2 min
TypeModel
TopicLlms · Gpus
  • AutoTrust AI Lab released GEV-26B-Decide, a Gemma-4-26B decision model with calibrated option probabilities.
  • System 1 answers in ~45 ms per decision; adaptive mode invokes reasoning only below 0.8 confidence.
  • Big gains: GPQA Diamond 42.9 to 78.6, BBH 75.0 to 92.0, MMLU-Pro 65.0 to 84.6.
  • Browser agent demo completes 95% of multi-step tasks at ~85 ms per click on one B200.
  • Supports yes/no, 0-5 score, and up to 256-way choice; 256K token context, image inputs.
  • Ships with vLLM server, POST /v1/decide endpoint, Apache-2.0 adapter and decision head.

GEV-26B-Decide turns Gemma 4 into a calibrated decision API

AutoTrust AI Lab has released GEV-26B-Decide, an open-weights model built on Gemma-4-26B-A4B-it. It maps typed questions to probability distributions over fixed options, then invokes step-by-step reasoning when the leading probability falls below a confidence threshold.

Applications can use those probabilities to route uncertain cases, select actions, and avoid parsing generated labels such as “A” or “B.” Calibration means that predictions assigned 80% confidence should be correct about 80% of the time.

  • Decision endpoint: Custom POST /v1/decide route alongside OpenAI-compatible routes
  • Inputs: Text and images, with contexts up to 256K tokens
  • Fast path: One forward pass, reported at about 45 ms
  • License: Apache-2.0 for the adapter, decision head, and calibration files
  • Base model: Covered separately by the Gemma 4 terms

One backbone, two inference paths

AutoTrust labels the two paths System 1 and System 2. System 1 reads a probability distribution from the model’s hidden state in one forward pass. System 2 runs the same loaded backbone in Gemma 4’s thinking mode, using the original state, question, and options as context. Both paths support text and image inputs through one vLLM engine.

System 1 accepts three decision schemas:

Kind Purpose Output space
noul Binary questions Yes or no
choice Option selection 2 to 256 choices
score Rating 0 to 5

Each request returns a probability for every available option. System 1 performs no token generation, which removes decoding latency and avoids extracting an answer from free-form text.

Uncertainty buys the extra compute

Adaptive inference uses the fast path unless its confidence gate triggers:

  1. Read the first distribution. System 1 produces p1 in one forward pass.
  2. Apply the confidence gate. System 2 runs when the highest probability in p1 falls below the default threshold of 0.8.
  3. Reason within a budget. The think_budget setting caps the number of thinking tokens.
  4. Read and combine. When the thinking channel closes, the model reads a second distribution, p2, and returns p = 0.5p1 + 0.5p2.

The 50/50 blend preserves information from the calibrated fast path because

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads