Open-Source Jeff Turns Tiny Models Into 22ms Local Decision APIs

Jeff packs zero-shot classification into 0.8B parameter models that decide in 22ms, trained on one workstation GPU in about two hours.

·
·
PRO
  • Jeff open-sources fine-tunes of Qwen3.5 (0.8B, 2B) and Gemma 4 E2B for zero-shot classification via GitHub
  • Overall benchmark score of 83.1 on the 2B model, essentially matching Jev's published 83.0
  • Latency of 22 ms per decision on RTX PRO 6000, 28 ms on Apple M4 Max via MLX
  • Trained on one workstation GPU in 2 to 3.5 hours using fully synthetic data from an open model
  • Voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in half an hour
  • Weights on Hugging Face under Apache 2.0; code under MIT

Jeff packages small models as a local decision API

Jeff is an open-source service that turns small language models into fast classifiers. The project fine-tunes Qwen3.5 0.8B and 2B models and Gemma 4 E2B to assign a calibrated probability to each supplied option in one forward pass. The endpoint generates no prose, so applications do not need to parse an LLM response.

On the hardware tested by the project, each decision takes about 22 ms on an RTX PRO 6000 and 28 ms through MLX on an Apple M4 Max. That performance gives developers a local function-like interface for routing, intent detection, moderation, grounding checks, and other tasks with clearly defined choices.

One endpoint, three output shapes

Jeff uses “zero-shot” to mean that its categories do not need to appear in the training data. A caller describes the current state and supplies options such as support queues, user intents, moderation labels, voice commands, or game moves.

The endpoint handles three output shapes and can evaluate several independent questions in one request, allowing related decisions to share the same input state.

  • choice selects among as many as 255 described options and returns a probability for each.
  • noul answers a yes-or-no question as a probability.
  • score selects a point on a scale defined by the caller.

Calibration comes from a fitted temperature applied after training. The adjustment changes the confidence distribution without adding another inference pass, with the goal of aligning predicted probabilities more closely with observed accuracy.

One request, two decisions

rust
curl -s localhost:8765/v1/systemone \
  -H 'content-type: application/json' \
  -d '{
    "model": "jeff-latest",
    "state": "Refund request: parcel arrived crushed, wants money back.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team?",
        "criteria": {
          "1": "Refunds",
          "2": "Damaged parcels",
          "3": "Account"
        }
      },
      "angry": {
        "type": "noul",
        "instructions": "Is the customer angry?"
      }
    }
  }'

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads