vLLM's Vela 2.0 Collapses Nine AI Safety Checks Into One Open Model

vLLM Semantic Router and KR Labs released four open models that answer safety, domain, PII and hallucination questions in one forward pass.

·
·
·
vLLM's Vela 2.0 Collapses Nine AI Safety Checks Into One Open Model
  • Vela 2.0 ships four Apache-2.0 routing models from 0.3B to 9B that answer safety, domain, PII and hallucination questions in one call.
  • Decoders are fine-tuned from Decision 2.0 Eos, Nox and Lux; the 0.3B encoder runs on CPU and exports to ONNX.
  • New span heads return labelled character spans for any label named in the request, trained with GLiNER2-style label resampling.
  • 9B hits 0.989 AUC on unseen prompt attacks and 0.774 F1 on RAGTruth hallucination, beating Vela 1.0 specialists and LettuceDetect v2.
  • Full router request takes 0.09 s on an A40 for the 0.3B and 0.60 to 0.71 s for the 9B, replacing nine separate models with one.
  • Models on Hugging Face, with the launch post covering the architecture and benchmarks.

Vela 2.0 folds nine LLM routing classifiers into one call

LLM routers commonly run several checks before selecting a backend model: attack detection, domain classification, personal-data detection, modality identification, safety review, and verification of claims against retrieved context. Vela 2.0 lets one checkpoint answer those questions in a single request, reducing the model loading and orchestration required by separate classifiers.

The vLLM Semantic Router team and KR Labs released four Apache-2.0 checkpoints in a public model collection. The 0.3B model is a CPU-oriented encoder that also exports to ONNX. The 0.8B, 4B, and 9B models are decoders fine-tuned from the Decision 2.0 Eos, Nox, and Lux bases, respectively.

Nine checks, one invocation

Vela 1.0 supplied nine single-purpose encoders for the router’s built-in signals: Safety, Hazard, Guard, Domain, PII, FactCheck, Halu, Feedback, and Modality. Each required a separate download, model load, and inference call. Adding a signal also required another dataset, model, and deployment path.

Vela 2.0 adopts the Decision model format, in which an application provides a question and its criteria at request time. The model supports four typed outputs:

  • Choice: selects one option from a supplied set.
  • Noul: returns a probability for a yes-or-no question.
  • Score: assigns one of several ordered levels.
  • Span: labels exact character ranges in the input text.

A single request can combine these output types. Domain and safety checks can return choices or probabilities, while PII detection and hallucination checks can identify the exact text that triggered a label. Applications may also define new labels through names and descriptions supplied with the request.

Questions stay isolated

Each decoder encodes the input text, called the state, once. Every question occupies a separate attention block that can read the state and its own tokens but cannot read another question’s block. In FP32, this isolation is exact: adding or removing a question leaves the other answers unchanged. Deployments can therefore vary checks by request while preserving cached raw scores, or logits, for stable questions.

Span extraction uses a word-by-label matrix of yes-or-no decisions. For each question, an FP32 bilinear layer and a small neural-network head score every word against every label description. The model applies a threshold, then merges adjacent words carrying the same label into character offsets. Two design choices adapt the process to a left-to-right causal decoder:

  1. A second copy of the target text appears after the label definitions, allowing each scored word to incorporate the complete text and every label.
  2. Each label representation is the mean of its full name and description block, allowing the model to distinguish labels introduced at request time.

On the team’s internal development data, these changes raised short-text PII F1 from 0.828 to 0.959. Long-document PII F1 increased from approximately 0.06 to 0.94.

The decoders contain two span heads with separate weights. A router head covers PII, unsupported claims, and toxic spans. A broad extraction head handles named entities, relations, and extractive evidence. A fixed rule assigns each question to a head. Because the broad head is trained last while the rest of the model remains frozen, router outputs stay bit-for-bit identical with the broad head loaded or absent.

Benchmark gains carry a tradeoff

The published evaluation reports that Vela 2.0 9B matches or exceeds each Vela 1.0 specialist on the specialist’s test set. The largest reported gains appear on attacks from unseen families and multilingual hate-speech checks.

Selected results reported for Vela 2.0 9B
Benchmark Comparison Vela 2.0 9B Metric
Unseen-family prompt attacks Vela 1.0 specialist: 0.792 0.989 AUC
Multilingual HateCheck Vela 1.0 specialist: 0.646 0.855 AUC
RAGTruth hallucination spans LettuceDetect v2: 3.1 points lower on identical rows 0.774 Example-F1
ACL-Verbatim evidence No GLiNER-family result exceeded 7.0 24.5 Word-F1

The ACL-Verbatim evidence benchmark was excluded from training, providing a test of the broad head’s ability to extract supporting text from unseen data.

Full-router latency reported on an Nvidia A40 GPU
Model Parameters Context Latency
Vela 2.0 0.3B 307M 8,192 tokens 0.09 seconds
Vela 2.0 0.8B 756M 16,384 tokens 0.13 seconds
Vela 2.0 4B 4.2B 16,384 tokens 0.40-0.49 seconds
Vela 2.0 9B 7.9B 16,384 tokens 0.60-0.71 seconds

Consolidating router tasks reduces performance on the Jev Decision Index, which contains 38 general decision tasks. The 9B model retains 89% of its Decision 2.0 Lux base score, the 0.8B model retains 79% of Eos, and the 4B model retains 74% of Nox. The release does not provide a retention figure for the 0.3B encoder. Decision 2.0 remains the higher-retention option for general decision workloads, while Vela 2.0 adds the safety, PII, hallucination, and extraction capabilities required by an LLM router.

Training proceeds in locked stages

Decoder training uses three stages, with each stage freezing the components completed earlier:

  1. Decision tuning: the Decision 2.0 base receives a 4,000-step fine-tune. A KL penalty limits changes from the frozen base model on replayed Decision 2.0 examples, helping preserve general decision behavior.
  2. Router spans: the router span head trains on PII, unsupported claims, and toxic text.
  3. Broad extraction: the broad head trains on a mix of 90% open extraction data and 10% replayed router-span data. The extraction set includes named entities, relations, SQuAD 2.0, HotpotQA, and Natural Questions evidence.
Largest components of the stage-one training mix
Data category Share of steps Coverage
Long-document PII 27.8% 17 entity types
Synthetic router decisions 23.1% 21 languages
Safety and prompt attacks 12.2% Includes AEGIS 2.0 and PolyGuardMix
Hallucination spans 11.3% LettuceDetect and RAGTruth training data

Label sets are resampled for every draw, with names anonymized, omitted, or paraphrased. This augmentation trains the model to interpret label descriptions supplied at inference time. Synthetic examples survive only when a blind relabeling pass agrees with the original label; the filter rejects approximately 29% of generated rows.

The 0.3B encoder follows a separate path. It trains from the Decision 1.0 Kai Choice trunk in one 101,000-step run, which took roughly 6.9 hours on a single AMD MI325X.

One schema across four sizes

All four checkpoints expose the same request and response schema. A team can prototype with the CPU-friendly 0.3B encoder, then move to a decoder by changing the checkpoint and runtime configuration while retaining the payload structure. The Semantic Router repository contains the surrounding router project.

python
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "vllm-sr/Vela-2.0-9B",
    trust_remote_code=True,
).to("cuda")

result = model.system_one(
    state={
        "request": (
            "Hi, I'm Tom Baker (tom.baker@example.com). "
            "What is the max daily paracetamol dose?"
        ),
        "source": "For adults, the max dose is 4 g in 24 hours.",
        "answer": "Adults can take up to 6 grams in 24 hours.",
    },
    questions={
        "pii": {
            "type": "span",
            "instructions": "Which spans are personal information?",
            "criteria": model.vela2_engine.cal["pii_schema"]["labels"],
            "over": "request",
        },
        "halu": {
            "type": "span",
            "instructions": (
                "Which spans of the answer are not supported by the context?"
            ),
            "criteria": {
                "unsupported": "a claim not supported by the context"
            },
        },
        "domain": {
            "type": "choice",
            "criteria": {
                "health": "...",
                "math": "...",
                "other": "...",
            },
            "over": "request",
        },
    },
)

The state object carries the text fields shared across checks. Each entry in questions defines an output type, instructions, and criteria; over can scope a check to one state field. The returned object contains one typed answer per question, with labels and character offsets for span results.

The trust_remote_code=True option executes Python code from the model repository. Production deployments can review that code and pin a repository revision before loading the checkpoint.

Each model also includes a FastAPI server, vela2_serve.py, which accepts the same payload at POST /v1/systemone. Automatic replacement of the nine Vela 1.0 encoders inside vLLM Semantic Router remains on the roadmap. Until that integration ships, developers must run a Vela 2.0 checkpoint or its server and connect the request payload to the router path themselves.

Trending
  • No trending articles

Comments

avatar

Next Reads