On Oct. 1, 2026, Cloudflare announced Clef and Clef-Flash as open-weight decision models: they return typed answers within a defined schema instead of generating open-ended prose. Cloudflare paired the release with hosted inference through Workers AI and weights under the Apache 2.0 license.
What a decision model returns
A decision model maps supplied information to an answer format set in advance. A developer can ask for a yes-or-no judgment, a choice from specified options, or a score; the model returns the decision in that structure rather than composing a free-form response.
That makes the output easier to use in a workflow that needs a defined result, such as a classification or a decision about what should happen next. The hosted API supports noul yes-or-no questions, choice, and score, and accepts 1–64 questions in a request.
Clef and Clef-Flash compared
The variants differ in model scale and backbone. Cloudflare describes Clef as a 27B model based on Qwen3.8-27B, and Clef-Flash as a 9B model based on Qwen3.5-9B. Both accept text, JSON, images, and video.
| Model | Parameters | Backbone | Inputs | Weights license |
| Clef | 27B | Qwen3.8-27B | Text, JSON, images, video | Apache 2.0 |
| Clef-Flash | 9B | Qwen3.5-9B | Text, JSON, images, video | Apache 2.0 |
Cloudflare lists a 65,536-token context window for hosted Clef. Its Workers AI service provides the hosted inference route; the released weights offer a separate route for local experimentation.
What Cloudflare’s evaluations report
Cloudflare’s published benchmark results vary by task. On BANKING77, it reports macro-F1 scores of 94.20 for Clef and 90.93 for Clef-Flash, compared with 79.74 for Jev. On When2Call, Jev leads with 80.97 accuracy, while Clef scores 72.37 and Clef-Flash 65.58. The figures describe different task-specific evaluations, not a single overall measure.
Cloudflare also described an internal website-classification workflow. It said Clef took 2.2 seconds to fetch, render, and classify a website; gpt-oss-120b took 4.7 seconds in the same workflow and returned two classifications. In Cloudflare’s Clef example, the model assigned 95% probability to fashion, 85% to ecommerce, and less than 1% to phishing.
Fine-tuning: hands-on first, self-service later
Cloudflare said it was offering fine-tuning through a hands-on service with its forward-deployed engineer (FDE) team. It described a self-service fine-tuning platform as a later development, rather than the service offered at the announcement.