Cloudflare releases Clef decision models, claims 39 ms median latency vs. 524 ms for TypeSafe's Jev
Cloudflare has released Clef and Clef-flash, two open-weight decision models that return probabilities over predefined answer options instead of generating text. The company claims median latencies of 39 ms and 209 ms, against just over 524 ms for TypeSafe AI's Jev. Both support text and images and are API-compatible with Jev.
Cloudflare has released Clef and Clef-flash, two decision models for AI agents. According to the company, Clef-flash returns results in a median of about 39 milliseconds and Clef in about 209 milliseconds, against just over 524 milliseconds for TypeSafe AI's Jev. Both models are available on Workers AI and on Hugging Face under the Apache-2.0 license.
What a decision model does
A decision model does not write prose. It answers several predefined questions about an input in one call and returns a probability for each answer option. For a customer support message, it can score urgency and pick the team that should handle the ticket. Downstream code then routes the ticket, escalates it, or hands it to a human.
Cloudflare positions these models between LLMs, whose outputs vary and can be slow, and traditional classifiers, which are fast but need retraining for each new category. The company says that agents can "programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed." Its API is fully compatible with Jev's, so customers can switch easily.
Specifications
- Clef: built on Qwen3.8-27B, according to Cloudflare
- Clef-flash: built on Qwen3.5-9B
- Inputs: text and images. Jev is text-only so far, per Cloudflare.
- Context window: 64,000 tokens, which Cloudflare says is twice Jev's
- License: Apache-2.0
- Pricing: not yet disclosed
- Training cutoff: not disclosed
Cloudflare says it leaves the base models unchanged and trains extra components on its own synthetic data. These components derive answer options and probabilities from the models' internal computations. It uses its own variant of Reinforcement Learning for Calibrated Decisions (RLCD), the method TypeSafe used for Jev. RLCD aims to make assigned probabilities match how often answers are actually correct.
Performance claims
Across 43 benchmarks, Cloudflare says both models are faster than all relevant competing decision models. All figures below are self-reported and not independently verified.
| Benchmark (accuracy) | Clef | Clef-flash | Jev |
|---|---|---|---|
| API Bank | 91.93 | 93.11 | 88.19 |
| When2Call | 72.37 | 65.58 | 80.97 |
| PhishNChips | 79.60 | 75.05 | 62.55 |
The results are mixed. Jev scores higher on When2Call, and Clef-flash edges out Clef on API Bank. Cloudflare says Clef leads on the Jev Decision Index and that Clef-flash nearly matches Jev's accuracy at a fraction of the latency.
In an internal test, Cloudflare's threat intelligence team used Clef to classify a website. It gave 95% probability of a fashion site, 85% of an online store, and under 1% of phishing. Fetching, rendering, and classifying took 2.2 seconds, versus 4.7 seconds for Cloudflare's fastest general-purpose LLM, which returned only two categories.
Fine-tuning service
Cloudflare is also launching a reinforcement learning service for customizing Clef. Forward deployed engineers will handle fine-tuning at first, with a self-service platform planned. Customers log requests through AI Gateway, evaluate them in containers that act as an RL sandbox, and deploy the tuned model on Workers AI through a new trainer component. Custom model serving uses technology from Replicate, which Cloudflare acquired in late 2025.
Cloudflare says it will use Clef internally to review abuse reports, sort support requests, and separate useful bots from harmful ones.
What this means
Decision models are becoming a distinct product category. TypeSafe AI introduced Jev in mid-September, and OpenAI followed in late September with a Decisions API built on GPT-6 Luna. Cloudflare is the third entrant in about three weeks, and the only one in this group that has released open weights.
The competitive angle is latency, price, and portability. Jev compatibility lowers switching costs, and Apache-2.0 weights let teams self-host. Cloudflare's own data shows Jev still ahead on at least one benchmark, so the accuracy lead is not uniform. Pricing is undisclosed, so cost comparisons must wait.
One caveat: calibrated probabilities do not make a model correct. They only keep it within predefined options. Removing humans from the loop depends on how well those probabilities hold up on real workloads, which independent testing has yet to show.
Related Articles
Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text
Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.
Amazon open-sources Strands Decider 2B, a small decision model built on a Qwen3.5-2B base
Amazon Web Services has released Strands Decider 2B, an open-source model that chooses among pre-decided options and returns a confidence score instead of generating text. It is inspired by TypeSafe's Jev and is small enough to run locally. Amazon says it briefly topped the Jevbench ranking for models of its size.
Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta
Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.
inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter
inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.
Comments
Loading...