Unsloth Turns a 0.8B Model Into a Fast Decision Engine at 78% Accuracy
Unsloth adds a Decision model trainer that turns small LLMs into calibrated classifiers, taking Qwen3.5-0.8B from 20.7% to 74.3% accuracy on 4GB VRAM.
- Unsloth adds Decision model training: fine-tune LLMs to output calibrated probabilities instead of text
- Qwen3.5-0.8B hits 74.3% aggregate accuracy across 3 benchmarks on just 4GB VRAM
- Uses a Clef-style head (same design as Cloudflare's Clef) trained with LoRA for one epoch
- Supports Qwen3.5, Llama 3.2, Gemma 4, with training times of 10 to 49 minutes
- Full GitHub repo and guide with notebooks available
- Deployable locally via a built-in Decision API compatible with Laya and Jev endpoints
Unsloth trains small LLMs to return typed decisions
Unsloth has added decision-model fine-tuning to its desktop app and open-source training stack, as described in its decision-model guide. Developers can adapt a pretrained language model to choose labels, answer yes-or-no questions, or score an ordered scale, with probabilities attached to the available options. The inference path returns structured values without autoregressive text generation, reducing output latency and eliminating malformed JSON.
A 0.8B model clears 70%
On Unsloth’s reported benchmarks, fine-tuning produced large accuracy gains for Qwen3.5-0.8B across three classification datasets:
| Dataset | Before | After | Gain |
|---|---|---|---|
| typed-decisions | 36% | 73% | 37 percentage points |
| BANKING77 | 7% | 74% | 67 percentage points |
| CLINC150 | 19% | 76% | 57 percentage points |
A separate 3,000-row held-out evaluation returned 78% accuracy, and the 0.8B model required 4GB of VRAM for training. These are vendor-reported results from Unsloth’s dataset mix and configuration. Production evaluation should also measure per-class accuracy, false-positive costs, calibration, latency, and performance on data from the intended deployment.
One pass, many typed answers
At inference time, the caller supplies source text followed by one or more typed questions and their permitted answers. The language-model backbone encodes that prompt once. A small classification head based on Cloudflare’s Clef design then scores the options from the model’s internal representations and returns a probability distribution for each question.
Autoregressive generation predicts output tokens sequentially and often requires a parser to validate the result. Unsloth’s decision path bypasses that decoding loop, which removes output-token latency and parsing failures. The backbone still processes the entire prompt, so adding questions and options increases sequence length and computation.
A support ticket can therefore produce several related decisions in one forward pass, such as the destination team, the detected intent, and whether the customer requested a refund. Each field receives its own selected value and probability distribution.

The recipe behind the gains
Unsloth’s published experiment used a Clef head and rank-64 LoRA adapters for one epoch. LoRA fine-tunes small, low-rank matrices while leaving most of the pretrained model unchanged, reducing the memory required for training. The starter configuration uses 4-bit base weights, rank-16 adapters, and a learning rate of 2e-4. Developers reproducing the benchmark should use its rank-64 configuration rather than assuming the starter defaults produce identical results.
The training mix combined typed-decisions with 12 classification and natural-language-inference sources: AG News, ARC, BANKING77, BoolQ, CLINC150, CommonsenseQA, MMLU, MNLI, prompt-injection examples, SNLI, SST-5, and WANLI. The 3,000-row test set drew from typed-decisions, BANKING77, and CLINC150. Unsloth says it decontaminated the test data against the training set.
What 4GB buys
Unsloth reports the following accuracy, peak VRAM, and training time across its tested models:
| Model | Test accuracy | VRAM | Training time |
|---|---|---|---|
| Qwen3.5-0.8B | 78% | 4GB | 42 minutes |
| Qwen3.5-2B | 81% | 8GB | 40 minutes |
| Llama 3.2 3B | 79% | 4.1GB | 30 minutes |
| Gemma 4 E4B | 77% | 14.4GB | 49 minutes |
| Laya, further fine-tuned | 77% | 2.5GB | 10 minutes |
The Laya result comes from further fine-tuning a pretrained decision model. Unsloth also reports that a shortened Qwen3.5-4B run with max_steps=60 reached 76% accuracy in 10 minutes on an Nvidia L4. Actual runtime and memory use will vary with hardware, sequence length, batch size, precision, and dataset shape.
From dataset to prediction
Unsloth Desktop can run the workflow without Python, while the open-source package exposes FastDecisionModel and DecisionTrainer for integration into existing pipelines.
- Select a base model such as
unsloth/Qwen3.5-4B. - Set Train as to Decision model.
- Choose the labeled dataset and start training.
Each supervised example needs source text and the expected answer for every trained question. The labels must match the question schema, including the permitted choices or scale values. The GitHub repository contains the trainer alongside Unsloth’s other fine-tuning tools.
In Python, predict accepts free-form input and a dictionary of typed questions:
answers = FastDecisionModel.predict(
model,
tokenizer,
"Hi, I was charged twice for invoice #4411. Please refund today.",
{
"team": {
"type": "choice",
"criteria": {
"billing": "invoices, payments, refunds",
"technical": "bugs, outages, errors",
"sales": "pricing, new plans",
},
},
"refund": {
"type": "bool",
"instructions": "Does the customer ask for a refund?",
},
},
)
print(answers["team"]["probabilities"])Measuring confidence with ECE
Each prediction includes the selected option and its probability distribution. The calibrate method compares those probabilities with observed correctness on held-out examples and reports accuracy, loss, and expected calibration error, or ECE. ECE groups predictions by confidence and measures the gap between average confidence and actual accuracy; lower values indicate closer agreement.
Calibration requires a separate, representative dataset. Probabilities calibrated on benchmark data can become unreliable when traffic, language, class frequency, or user behavior changes. Deployed systems should monitor both accuracy and calibration, then recalibrate or retrain when those measures drift.
Choose tasks with finite answers
Strong candidates
- Ticket routing and intent classification
- Content-moderation triage
- Prompt-injection screening
- Document tagging and workflow branching
- Risk scores with explicit thresholds
- Several related decisions derived from the same input
Constraints to test
- Every answer must fit a predefined choice, Boolean field, or scale.
- More questions and options increase prompt length and inference cost.
- Calibration can degrade under distribution shift.
- Aggregate accuracy can conceal weak performance on rare or costly classes.
- A conventional encoder classifier may use less memory and run faster, so it remains a useful baseline.
- Open-ended writing and tasks requiring generated explanations need a generative output path.
Teams using a general-purpose model API solely to classify text can now evaluate a local 0.8B or 3B decision model with structured outputs. The useful comparison covers task-level accuracy, calibration, peak memory, p95 latency, operating cost, and maintenance effort across the decision model, a conventional classifier, and the existing API.