model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

TL;DR

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

3 min read
0

Cloudflare has released Clef, a 27B-parameter multimodal model that converts a state and a schema of typed questions into probabilities for each allowed answer option. It does this in a single forward pass, with no free-form text generation and no output parsing. The weights are on Hugging Face under the Apache-2.0 license.

Architecture and interface

According to the model card, Clef is post-trained from Qwen/Qwen3.8-27B and keeps that model's vision encoder. The backbone is stored as standard sharded safetensors. A small transformer "joint schema head" reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly. The output is one logit per allowed option, and a softmax per question yields probabilities.

Input state can be a string, JSON, images, or video frames. Three question types are supported:

  • noul: true/false, returning the probability of true
  • choice: named options, returning the choice, confidence, and probabilities
  • score: ordered options, returning an expected score, confidence, legend, and probabilities

The encode_record function defaults to max_length of 16,384 tokens. Cloudflare does not state a model context window. The card says the Clef API is fully compatible with Jev and SystemOne, and a systemone function mirrors the POST /v1/systemone request and response format. A smaller, faster variant, Clef-Flash, is also available. Its parameter count is not given in the source.

Cloudflare tested the code with torch 2.11 and transformers 5.10.2 on a single H200. Pricing is not yet disclosed, since the weights are self-hostable.

Benchmarks

All figures come from Cloudflare's own run of its Decision Index 0.2.1 suite and have not been independently verified. Comparison columns are Clef-Flash, Jev, DiffusionGemma Jev, Kev 9B, and Laya.

Benchmark Clef Jev
BFCL (case exact accuracy) 98.5 95.8
BANKING77 (macro-F1) 94.2 79.7
CLINC150+OOS (macro-F1) 97.4 89.3
MMLU 90.3 91.7
MMLU-Pro 65.9 82.7
GPQA Diamond 48.0 78.3
BBH 73.7 92.9
GSM8K 80.8 79.9
ForecastBench (Brier, lower is better) 13.9 17.4
Median latency (ms) 209.3 524.1

Clef trails Jev on general-knowledge and reasoning tests, including GPQA Diamond, MMLU-Pro, and BBH. Clef-Flash posts a median latency of 38.8 ms, and its ForecastBench Brier score of 10.6 is better than Clef's. Clef-Flash also beats Clef on several tasks, including the home appliance simulator (97.7 vs 83.0) and CLadder (97.7 vs 94.0). Clef is clearly ahead of Clef-Flash on CLINC150+OOS (97.4 vs 66.8) and RAGTruth hallucination F1 (79.4 vs 35.6).

Workflow evals

On four business workflows from Typesafe Evals, scored against consensus reference labels, Clef edges Jev on three:

  • Invoice processing: 64.7 exact actions and 86.2 primary action, vs 61.8 and 83.1 for Jev
  • Customer service: 76.3 exact actions, vs 76.0 for Jev and 77.0 for Clef-Flash
  • Security incidents: 62.9 exact actions, vs 61.7 for Jev
  • Agent trace observability: 68.5 primary action, behind Jev's 71.6 and Clef-Flash's 69.8

What this means

Clef targets classification, routing, triage, and tool-selection tasks where a downstream system needs calibrated probabilities rather than text that must be parsed. Returning scores for every option in one pass removes parsing failures and makes latency predictable. Cloudflare's reported 209 ms median is less than half of Jev's 524 ms.

The benchmark table shows the trade-off. Clef is strong on structured decision benchmarks but weak on broad reasoning relative to Jev. It is a specialized decision component, not a general assistant. The Apache-2.0 license and Jev/SystemOne API compatibility make it easy to try as a drop-in. But every number here is Cloudflare's own, on its own suite, so teams should validate on their own schemas before relying on it. Hardware for the latency figures is unspecified.

Related Articles

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

model release

Google unveils Gemini 4 Argon at $2/$10 per 1M tokens, but access is limited to select users

Google unveiled Gemini 4 Argon on Wednesday with introductory pricing of $2 per 1M input tokens and $10 per 1M output tokens, matching OpenAI's discounted GPT-6.1 Sol. Access is restricted to select cybersecurity defenders and enterprise cloud customers, and Google says published rates will double later.

model release

Ideogram 4.5 launches with native 2K output and four tiers from $0.008 to $0.22 per image

Ideogram has released Ideogram 4.5, an image model it claims edits only the area a user specifies and leaves the rest untouched. It offers four quality tiers from 0.8 to 22 cents per image, all at native 2K resolution, via the Ideogram platform and API. An open-weight release is promised but not yet dated.

model release

Amazon open-sources Strands Decider 2B, a small decision model built on a Qwen3.5-2B base

Amazon Web Services has released Strands Decider 2B, an open-source model that chooses among pre-decided options and returns a confidence score instead of generating text. It is inspired by TypeSafe's Jev and is small enough to run locally. Amazon says it briefly topped the Jevbench ranking for models of its size.

Comments

Loading...

Cloudflare Clef: 27B Multimodal Decision Model, Apache-2.0 | TPS