AI NewsModels & agentsAnnouncement
Cloudflare AI Gateway now picks a model per request and says its own tests came in 30 percent cheaper than Opus and Sol
Cloudflare opened its Auto Router in public beta on 30 September, routing each AI Gateway request to a smaller or larger model based on a classifier and reporting up to 30 percent cost savings against the strongest available models in Cloudflare's own tests.

Image: Cloudflare
Why it mattersTeams that let every engineer pick a model per prompt end up paying the top-tier price for tasks that do not need it, and a router that decides per request removes that choice from the individual without blocking the strong model when the work needs it.
A team that hands out API keys to every engineer ends up paying the top-tier price for jobs a cheaper model could answer. Cloudflare said on 30 September its AI Gateway now has an Auto Router in public beta that picks a model per request based on a classifier running on Workers AI, and reports its own results as 30 percent cheaper than routing everything through OpenAI's GPT-6 Sol or Anthropic's Claude Opus 5.5.
The setup is small. Point a request at the model name cloudflare/auto through AI Gateway. The gateway reads a short slice of the recent conversation, runs it through a multi-head classifier that assigns probabilities across 14 task categories, rates the request on complexity, ambiguity, stakes and dependence on earlier context, then picks a model from the pool that survives the gateway's own filters (format support, credentials, spend limits, provider health).
The benchmark, from Cloudflare, against Cloudflare
Cloudflare says it uses Auto Router itself inside its OpenCode deployment and its Cloudflare OS agent harness. The published comparison is Cloudflare's own general-knowledge-work benchmark of 97 tasks, three samples per task per model, covering email, calendars, Slack, files, travel and finance.
| Model | Successful trials | Success rate | Total cost | Cost per success |
|---|---|---|---|---|
| cloudflare/auto | 252/291 | 86.6% | $2.10 | $0.0084 |
| Anthropic Claude Opus 5.5 | 281/291 | 96.6% | $5.91 | $0.0210 |
| OpenAI GPT-6 Sol | 245/291 | 84.2% | $2.64 | $0.0108 |
Cloudflare reports 95 percent confidence intervals of +6.2 or -6.9 percentage points on Auto Router, +2.7 or -3.8 on Opus, and +6.5 or -6.9 on Sol, from 10,000 task-level bootstrap resamples. In dollar terms, Cloudflare's number is 80 percent of Sol and 35 percent of Opus. Opus still ranked highest on success at 96.6 percent, which the post is straight about: Auto Router does not beat the strongest model on a task, it declines to pay for the strongest model on tasks that a cheaper one clears.
What the tables leave out
Every figure in the table is Cloudflare benchmarking its own routing product against two competitors on its own benchmark, so read it as the vendor's number. The post does not publish per-category results, does not name every model in the routing pool, and does not describe how a task's difficulty was labelled to check the router's choice. It also acknowledges that a cheaper price per token can end up more expensive per outcome, because a weaker model uses more tokens per solve, and says the router optimises for predicted trajectory cost rather than dollars per million tokens.
The rules that let it work
The router runs on Workers AI on GPUs across Cloudflare's edge, so the classifier itself sits in the request path. Cloudflare says it filters out unhealthy upstream providers, brings them back after outages, and applies the gateway's existing controls: credentials, billing configuration, access policies and spend limits. Requests that need a specific model still get it: this is the routing default, not a lock.
For a team that has already put AI Gateway between its engineers and the providers (for spend visibility, per-user analytics, budget caps), Auto Router is one more setting on the same product. It moves the "does this task need Opus?" decision away from the individual writing the prompt, without blocking Opus for the work that needs it. Whether the 30 percent survives a team's own workload is a small experiment, run against your own traffic in front of the change, before trusting the Cloudflare benchmark.
Source
- Cut your AI spend with AI Gateway's Auto Router, Cloudflare blog, 30 September 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


