AI NewsModels & agentsAnnouncement
Cloudflare releases Clef-omni, drops Clef-flash price, speeds up Clef
Cloudflare has added audio and video input to its open-weight Clef decision model with Clef-omni, cut Clef-flash from 0.09 to 0.038 dollars per million input tokens, and sped up Clef by 1.7 to 2.0 times with no new weights.

Image: Cloudflare
Why it mattersA team that was wiring up transcription and vision pipelines so an LLM could decide between options can now send the audio and video to one scoring call that returns calibrated probabilities.
Deciding whether a 21-second clip shows a running fan or a stopped one used to mean stringing together speech-to-text, image classification and a language model, then hoping the final step read the pieces the same way a person would. Cloudflare's new Clef-omni takes the audio, the video frames and a text prompt in one request and returns a calibrated score for each option it is given.
Clef-omni joins Clef and Clef-flash in the Clef family of open-weight decision models Cloudflare launched earlier this month. The three come with a price cut on Clef-flash, from 0.09 to 0.038 dollars per million input tokens, and a serving-side speedup that makes Clef itself 1.7 to 2.0 times faster. All three models speak the same Jev-compatible API, so client code written for TypeSafe's Jev needs no changes.
What Clef-omni actually scores
The new model takes audio (wav or mp3), video (mp4 or webm), images and text in one API call. Cloudflare says text-only decisions return in about 130 ms at the median, image inputs in about 150 ms, audio clips in a few hundred milliseconds, and a 21-second video clip with sound in about 1.5 seconds. Clef-omni is built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts backbone, with the text-to-speech output components removed and low-rank adapters trained on top. The weights are open on Hugging Face. Pricing is 0.15 dollars per million input tokens.
On the benchmarks Cloudflare publishes, Clef-omni scores 98.2 on BFCL case-exact, 92.7 on API-Bank, 94.8 on BANKING77 macro-F1, and 97.7 on CLINC150+OOS macro-F1. On the Home appliances benchmark Clef-omni scores 69.3, where Clef-flash scores 97.73 and Jev scores 52.27. These are Cloudflare's own numbers, measured on its own evaluations, and the post reports them against Clef, Clef-flash and Jev in the same table.
The Clef-flash price cut comes with one trade-off
Clef-flash now sits at 0.038 dollars per million input tokens, which Cloudflare describes as cheaper than Jev. The trade-off is on the hosted version's context window, which drops from 64k tokens to 24k. Cloudflare says 0.24% of requests have been exceeding 24k tokens, and recommends teams with larger context needs switch to Clef, which keeps its 64k window and its 0.24-dollar price. The weights on Hugging Face still support a 256k context window for anyone who self-hosts.
The Clef speedups are infrastructure changes with no new weights. Cloudflare moved the hosted serving to SGLang and worked with the SGLang team to merge support for Clef, which lands in SGLang 0.5.22. On input of about 800 tokens the median response moves from 262 ms to 152 ms, and on about 3,400 tokens from 616 ms to 305 ms. The same optimisations are reflected in a new launch command for anyone self-hosting Clef through SGLang.
Cloudflare names four internal uses of Clef so far: closing spam issues on the public GitHub docs repository, moderating plugin libraries for phishing in its EmDash content system, scanning data for personally identifiable information, and detecting malicious domains in its threat intelligence work. Each is a classification job that used to need a trained small model or a full LLM call, and each pulls a calibrated probability out of a single scoring request.
Source
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.

