Cloudflare's Clef Beats Rivals With 38ms Typed Decisions for AI Agents

Cloudflare released Clef and Clef-flash, open-source multimodal decision models that return calibrated probabilities over typed schemas instead of free-form text.

·
·
Cloudflare's Clef Beats Rivals With 38ms Typed Decisions for AI AgentsPRO
  • Cloudflare released Clef and Clef-flash, Apache 2.0 decision models post-trained from Qwen3.8-27B and Qwen3.5-9B.
  • Models return calibrated probabilities over typed schemas in a single non-autoregressive forward pass, with no text parsing.
  • Multimodal input and 64k context window beat Jev's text-only 32k, with vision encoder inherited from Qwen.
  • Median latency 209 ms for Clef and 38.8 ms for Clef-flash, versus 524 ms for Jev on Decision Index.
  • Hosted on Workers AI with a Jev-compatible API, plus local support via vLLM, SGLang, Transformers.
  • Companion RL fine-tuning service lets customers specialize Clef using AI Gateway traffic and Workers AI deployment.

Cloudflare’s Clef targets fast, typed decisions for AI agents

Cloudflare has released Clef models, a family of open-weight models built to select among predefined options and return calibrated probabilities. Clef and the smaller Clef-flash use the Apache 2.0 license and are available through Hugging Face, Workers AI, and local inference frameworks.

Agentic applications repeatedly make bounded decisions: choosing a tool, routing a ticket, assigning severity, or deciding whether to escalate. General-purpose language models can handle those tasks through structured generation, but they still generate tokens and require output validation. Clef scores the allowed answers directly, which gives developers typed results without parsing free-form text.

Cloudflare positions Clef against Typesafe AI’s Jev System One, another model designed for decisions inside agent pipelines. Cloudflare reports lower median latency on its Decision Index 0.2.1 run, stronger results on several classification benchmarks, vision support, and a context window of up to 64,000 tokens.

Decisions without token generation

A decision model accepts a description of the current state and a schema containing typed questions. It returns a probability distribution over the allowed answers for each question. Applications can then apply thresholds, trigger human review, or pass the selected values to another workflow step.

Clef supports three question types:

  • noul for binary true-or-false decisions
  • choice for named options with descriptions
  • score for ordered categories such as severity levels

A request can combine all three types:

json
{
  "state": "Checkout has been failing for every customer for the last hour.",
  "questions": {
    "urgent": {
      "type": "noul",
      "instructions": "Is this urgent?"
    },
    "team": {
      "type": "choice",
      "criteria": {
        "billing": "Payments and invoices",
        "technical": "Outages and errors"
      }
    },
    "severity": {
      "type": "score",
      "criteria": ["None", "Minor", "Major", "Critical"]
    }
  }
}

Clef evaluates the questions together and returns probabilities for every permitted option in one model pass. The caller retains control over the final policy, including minimum confidence, fallback behavior, and escalation rules.

Labels that change with each request

A conventional supervised classifier usually has a fixed output layer, so adding or renaming labels can require retraining. Clef receives the question schema as part of each request, allowing one deployment to handle changing labels and multiple decision fields.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads