AI NewsModels & agentsAnnouncement

Cloudflare released Clef, an open-source decision model that is API-compatible with TypeSafe's Jev and beats it on eight of ten benchmarks in its own tests

Cloudflare released Clef and Clef-flash, two open-source decision models that return typed probabilities instead of text. The models are hosted on Workers AI, released on Hugging Face under Apache 2.0, and are a drop-in replacement for TypeSafe AI's Jev API.

AI News

Editorial2 min read

LinkedInX
Cloudflare Clef decision models announcement header

Image: Cloudflare

Why it mattersA team already using Jev for classification can now swap the endpoint and read Cloudflare's own benchmark table to decide which model fits each decision, with no code change beyond the URL.

Cloudflare announced Clef and Clef-flash today, two decision models that return typed probabilities instead of text and are a drop-in replacement for TypeSafe AI's Jev API. The models are hosted on Workers AI and released on Hugging Face under Apache 2.0, so a team can call the hosted endpoint or run the weights on its own machine.

Cloudflare says Clef currently leads the Jev Decision Index. The post includes a benchmark table against Jev, DiffusionGemma, Jev Kev-9B and Laya, and the full scores are on Cloudflare's live demo site.

What the model does and how it differs from Jev

A decision model takes an input and a set of typed questions, and returns probabilities for each answer. Cloudflare's example is a support message that asks three questions in one call: is this urgent, which team should handle it, and how severe is the impact. The output is structured so a Worker can route the ticket or escalate without reading free text.

Cloudflare says Clef has a vision encoder, so it can classify images as well as text, which Jev does not do today. The context window is 64,000 tokens against Jev's 32,000. The architecture does a prefill-only pass through a Qwen backbone, then scores schema choices in parallel, so the decision step is non-autoregressive and no text is generated token by token.

The numbers Cloudflare published

Cloudflare's table reports Clef at 98.47 against Jev's 95.75 on BFCL case-exact, 91.93 against 88.19 on API-Bank accuracy, and 94.20 against 79.74 on BANKING77 macro-F1. Jev wins When2Call accuracy at 80.97 against 72.37 and BRIGHT nDCG at 47.52 against 45.91. On a four-workflow suite from TypeSafe's own evals, Cloudflare says Clef beat Jev in three of the four.

On latency, Cloudflare reports median response times of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash against 524.1 for Jev. Every figure above is Cloudflare's own measurement.

The RL fine-tuning service

Alongside the models, Cloudflare opened a reinforcement learning service that fine-tunes Clef for a specific workload. The first form is hands-on work with its forward-deployed engineer team. A self-serve platform where customers capture data, fine-tune and redeploy comes later. Cloudflare's internal uses, named in the post, are classifying submitted domains for the Threat Intelligence team, triaging Cloudflare Support requests, and sorting good crawlers from bad ones in Bot Management.

A team evaluating a classifier for a production path now has a second vendor to compare against Jev that is open source, API-compatible with the first, and benchmarked by its author against the first. The benchmark table is the vendor's, so the honest next step is to replay those evals on the team's own data before deciding which model to put in a live production path. For a team that cares about the licence, Apache 2.0 lets Clef run on self-hosted hardware without a Cloudflare account at all.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX