Cloudflare releases open-weight Clef decision models and RL fine-tuning

Cloudflare has released two models it trained itself, Clef and Clef-flash. They are what the company calls decision models: small, fast models that return bounded, structured answers with probabilities, meant to be dropped into a workflow wherever a decision is needed. Cloudflare frames the release as a response to recent buzz around Typesafe AI's Jev System One model, which it credits with introducing the decision-model concept. Such models can work over any set of inputs without being retrained every time a new category appears, in contrast with LLMs, which are largely non-deterministic but open-ended enough to reason and generate text and tool calls.
The models are hosted on Workers AI and are fully Jev-API compatible, so people can try them as a swap-in. Cloudflare is also open-sourcing them on Hugging Face under an Apache 2.0 license for local use. It says Clef is currently the leader on the Jev Decision Index, with full results on a live benchmark demo site. That is Cloudflare's own evaluation.
The post explains a decision model with a support-ticket example: pass in a customer message and ask whether it is urgent and which team should handle it. The model returns typed answers with probabilities, which code can use to route the ticket, escalate, or defer to a human. Cloudflare's own example comes from its Threat Intelligence team, which has been testing Clef to classify website domains. Given a domain (with Browser Run), Clef returns category probabilities, for example a 95% chance of a fashion website, 85% ecommerce and under 1% phishing. In that workflow Clef took 2.2 seconds to fetch, render and classify the site. Cloudflare's fastest general LLM, gpt-oss-120b, took 4.7 seconds in the same workflow and returned only two classifications. Cloudflare calls this a 2x saving in latency and results. It is a single example, not a benchmark average.
Cloudflare names three differences from Jev. Clef has a vision encoder, so it can classify images, while Jev only does text classification today. Clef has a 64k context window against Jev's 32k. And it scores competitively on quality benchmarks drawn from the Jev Decision Index. On Typesafe's own eval suite, Cloudflare says the Clef models beat Jev in 3 out of 4 areas, and that Clef-flash performs exceptionally well given how much faster it is. Across 43 eval benchmarks, the Clef models beat the other decision models on latency, except Laya, which Cloudflare describes as very fast but weaker on quality. The benchmark tables themselves are not reproduced in the text. Because the models run on Cloudflare's edge GPUs, the company says network latency is low enough to put Clef in the hot path for agents and pair it with an LLM on Workers AI that takes the action. The larger Clef is the precision model; Clef-flash targets latency-critical decisions.
On training, Cloudflare says it adapted DiffusionGemma in an earlier demo to output deterministic probabilities by exposing LLM logprobs, building on independent research by Matt Mastracci. Clef uses a different backbone: Qwen. The post names Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, both frozen. At inference Clef runs a prefill-only pass and scores the valid schema choices in parallel. The decision step is non-autoregressive, so no intermediate text is generated token by token, which Cloudflare credits for the speed advantage. A two-stage attention routing process lets each valid choice extract relevant context, and lets field parameters cross-attend with each other and with the original payload before scoring. The routing head was optimized jointly with rank-256 low-rank adapters. Post-training used label-smoothed cross-entropy for valid schema outputs plus a Brier loss to refine probability calibration, on internal synthetic datasets that permute field orders, prompts and schema structures. Cloudflare also developed Reinforcement Learning for Calibrated Decisions (RLCD) as a secondary target: it gives partial credit to adjacent ordinal choices, rewards fully precise record outputs, and applies a reference penalty to limit distribution shift.
The third announcement is an RL fine-tuning product for Clef. Internal Cloudflare teams want fine-tuned classifiers for Trust & Safety submissions, triage of Support requests, and deciding whether a crawler is a good or bad bot in its Bot products. Cloudflare says it has more than 15 years of network data across different domains and many years of labelled decisions to train on. The service launches as a hands-on engagement with its forward-deployed engineer (FDE) team; the company says it will learn from that to build a self-serve platform where customers capture data, fine-tune and redeploy the model, all on Cloudflare. The building blocks are AI Gateway (collects AI traffic into a dataset), Workers AI (generates rollouts against the base Clef model), Containers (an RL sandbox for scoring and replaying agent actions), a new Trainer component (updates the fine-tuned weights), and Workers AI with BYO Model (redeploys the result). The post says Cloudflare does not read, store or train on requests or responses, unless the customer uses the fine-tuning product.
Key facts
- Cloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and Jev-API compatible.
- Cloudflare says Clef leads the Jev Decision Index; it also has a vision encoder (Jev is text-only today) and a 64k context window versus Jev's 32k.
- In one Threat Intelligence domain-classification example, Clef took 2.2s versus 4.7s for gpt-oss-120b, which returned only two classifications.
- Clef uses a non-autoregressive decision step on a frozen Qwen backbone (Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash) with rank-256 adapters and an RL method called RLCD.
- An RL fine-tuning service for Clef starts as a hands-on offering with Cloudflare's forward-deployed engineers, with a self-serve platform planned later.
Why it matters
Cloudflare is entering the new decision-model category with an open-weight model. Decision models return bounded, typed answers with probabilities rather than free text, so code can route a ticket, escalate or hand off to a human without a person in the loop and without an LLM generating tokens. Clef adds image input and a longer context than Jev, and Cloudflare says it leads Jev's own index. The RL fine-tuning service also signals that Cloudflare wants to be where customers train and redeploy small specialised models.
Who it affects
Developers building agentic workflows who need quick, consistent classification steps, such as support routing, trust and safety triage, bot detection or domain categorisation. Teams already using Jev can try Clef, since it is Jev-API compatible. Customers with large sets of labelled decisions are the target for the fine-tuning service. Cloudflare's own Threat Intelligence, Trust & Safety, Support and Bot teams are named as internal users.
How to use it
Call the hosted model through the Workers AI endpoint (@cf/cloudflare/clef). The request carries a "state" text and a set of "questions", each with a type (the example shows noul, choice and score), instructions and criteria; outputs come back typed with probabilities. Developer documentation and a Hugging Face repo are available, and the weights can be run locally under Apache 2.0. Clef is the precision model; Clef-flash suits latency-critical decisions. For fine-tuning, the first route is working with Cloudflare's FDE team. The post gives no pricing for hosting or fine-tuning and no timeline for the self-serve platform.
How solid is it
The performance claims are Cloudflare's own: the leading position on the Jev Decision Index, the 3-of-4 result on Typesafe's eval suite and the latency wins across 43 benchmarks, and the post mentions no independent verification. The tables are not reproduced in the text, so there are no per-benchmark scores or figures for Clef-flash, Laya or Jev. The 2.2s versus 4.7s comparison is one Threat Intelligence example, not an average. The post's description of the RL platform stops partway through, so details after the list of building blocks are unknown.
Risks and caveats
Cloudflare itself notes that fine-tuning may give up some general-purpose performance in exchange for higher accuracy in one domain. The privacy promise (no reading, storing or training on requests and responses) does not apply if a customer uses the fine-tuning product. Some pieces of the RL platform are described as work in progress, and the self-serve version is a later step. Cloudflare notes that Laya is faster than Clef but trades off quality. The post gives no parameter counts beyond the backbone names.
“A decision model makes classifications to help agents decide how to act, based on certain probabilities.”
— Cloudflare blog post introducing Clef