AutoTrust's JEV-27B-VL Beats GPT-4o and Scores Images in One Pass
A vision-capable 27B decision model that returns calibrated probabilities in one forward pass, matching collaborative filtering with zero training data.
- JEV-27B-VL adds vision to the JEV decision model, built on Qwen3.8-27B under Apache-2.0.
- System 1 returns calibrated probabilities for yes/no, 0-5 ratings, or 2-256 choices in one forward pass.
- Zero-shot video recommendation from covers matches collaborative filtering trained on 59,045 users (AUC 0.727 vs 0.728).
- Tops VL-RewardBench at 78.3% and Plan-RewardBench at 73.2%, beating GPT-5 and Gemini-3-Flash on agent judging.
- Handles 256K-token prompts and ships a
POST /v1/decideendpoint on top of vLLM. - Robot pick-and-place in MuJoCo at 75% success, browser automation at 95% over 60 multi-step tasks.
JEV-27B-VL Scores Text and Images in One Pass
JEV-27B-VL adds a typed decision interface to a vision-language model. Given text, images, or both, it can answer yes-or-no questions, choose among 2 to 256 options, or assign a 0-to-5 rating. Each request returns a probability for every option in one forward pass.
The Apache-2.0 release combines a Qwen3.8-27B backbone with a LoRA adapter and decision head. Its Hugging Face page has recorded more than one million downloads. The decision head scores all options together, avoiding token-by-token label generation and making the model suitable for ranking, routing, evaluation, and control loops.
A calibrated probability is intended to track real-world frequency. Among decisions assigned a probability near 0.8, roughly 80% should prove correct when calibration holds. That property supports confidence thresholds and escalation policies, although the authors have not systematically measured calibration on image tasks.
Two paths through one model
| Mode | Behavior | Best fit |
|---|---|---|
| System 1 | Applies the adapter and decision head to text and images, supports prompts up to 256K tokens, and returns probabilities for typed options. | Classification, ranking, judging, routing, and repeated low-latency decisions. |
| System 2 | Uses the Qwen3.8-27B backbone for free-form generation, with optional step-by-step reasoning and image input. | Open-ended tasks that require explanation, synthesis, or additional reasoning. |
The vision release retains the text decision capabilities of JEV-27B while adding screenshots, photographs, and video cover images as inputs. Developers can route routine decisions through System 1 and send uncertain or open-ended cases to System 2.
Cold-start ranking from six covers
The short-video demonstration gives System 1 the cover images from a user’s five most recently watched videos plus one candidate cover. The model returns the estimated probability that the user will watch the candidate. This setup requires recent user context, but it does not use population-level interaction logs, learned item embeddings, or recommender-specific training.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.