最新模型
Cloudflare Open-Sources Clef Decision Models, Producing Option Probabilities in a Single Forward Pass
Clef and Clef-flash turn agent classification and routing into direct scoring of specified options, and release their weights under Apache 2.0. Official tests show latency advantages, but task accuracy and probability calibration still need to be validated against deployment data.

On October 1, Cloudflare released Clef and Clef-flash, making them available as hosted services on Workers AI and publishing the models. The technical focus is enabling agents to make decisions from a predefined set of answers: for example, deciding which team should handle a customer service message or whether it requires immediate attention. Code can then use the option probabilities to determine what happens next. Official announcement
The two models are built on Qwen3.8-27B and Qwen3.5-9B, respectively. The Clef model card explains that a new joint structured decision head reads the backbone’s final-layer hidden states, routes input evidence to each question, and jointly calculates scores for all valid options. A softmax converts the scores for each question into probabilities. The entire process takes only one forward pass, avoiding answer generation token by token and subsequent text parsing. Clef model card
The public interface accepts a state describing the scenario and a set of questions. It supports named options, ordered scoring, and true-or-false judgments, and can also take images or video. The systemone function included with Clef-flash accepts Jev/SystemOne request formats and returns answers, confidence values, and probabilities, giving existing decision workflows a concrete way to switch over. The public files include the weights, decision head, and encoding code, under the Apache 2.0 license. Clef-flash model card
Performance should be assessed task by task. In Cloudflare’s internal Decision Index tests, the reported median request latency was 209.3 milliseconds for Clef and 38.8 milliseconds for Flash. However, Flash scored 66.8% macro-F1 on CLINC150+OOS, below Clef’s 97.4%. These are publisher-run tests; they do not establish that Chinese-language classification, different hardware, or production services will see the same benefits. Test results
Engineering teams could first evaluate nodes with clearly defined answer sets, such as tool selection and ticket routing. During deployment, they should validate using the custom decision head examples in the model card and measure performance on Chinese-language data, unknown categories, and error costs. The output probabilities also need calibration checks to set thresholds for handing cases to a human. Constraining the output format does not guarantee correct judgments.