Jev-Omni Scores Decisions Across Text, Images, Audio, and Video in Milliseconds
A 12B multimodal classifier fine-tuned from Gemma 4 that returns calibrated probabilities across text, image, audio and video inputs.
- Jev-Omni is a 12B multimodal decision classifier built on Gemma 4 12B IT, Apache-2.0 licensed.
- Handles text, image, audio and video in one model, returning calibrated probabilities over supplied options.
- Scores 87.57% on DecisionBench Medium, 63.10% MMAU, 53.10% MVBench; ECE of 0.0400.
- Warm H200 latency: 83ms text, 26ms image, 31ms audio, 504ms 16-frame video.
- Best at 20 or fewer options; needs ~50GB FP32, inference in BF16 on CUDA GPU.
- Try it free on the hosted Space; independent of TypeSafe AI's Jev.
Jev-Omni scores fixed choices across four modalities
Jev-Omni is a 12-billion-parameter open-weight classifier that accepts text, images, audio, or video and assigns a probability to each supplied answer. A request contains a state, meaning the context being evaluated, along with a question and candidate options. The fixed output contract gives applications structured decisions without parsing generated prose.
The model was fine-tuned from Gemma 4 12B IT on 30,000 questions. Its author claims it is the first downloadable checkpoint to support all four modalities through one decision interface, a combination that many open-weight alternatives lack. The repository lists an Apache-2.0 license; adopters should also review the terms attached to the upstream Gemma components downloaded by the loader.
Fixed choices, usable probabilities
Generative multimodal models emit token sequences that structured pipelines must validate and parse. Jev-Omni directly scores the options supplied with each question, reducing the need for regular expressions, constrained decoding, or a separate judge model.
The typed-decision interface supports yes/no, multiple-choice, and score questions. On the medium split of DecisionBench, the model card reports an expected calibration error of 0.0400 across 10 confidence bins. Expected calibration error measures the weighted gap between predicted confidence and observed accuracy, so 0.0400 represents an average gap of about four percentage points on that test split. Calibration can change with new data and deployment conditions.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.