AI NewsModels & agentsAnnouncement

Cloudflare AI Gateway now picks a model per request and says its own tests came in 30 percent cheaper than Opus and Sol

Cloudflare opened its Auto Router in public beta on 30 September, routing each AI Gateway request to a smaller or larger model based on a classifier and reporting up to 30 percent cost savings against the strongest available models in Cloudflare's own tests.

AI News

Editorial3 min read

LinkedInX
Cloudflare blog card for the Auto Router public beta announcement

Image: Cloudflare

Why it mattersTeams that let every engineer pick a model per prompt end up paying the top-tier price for tasks that do not need it, and a router that decides per request removes that choice from the individual without blocking the strong model when the work needs it.

A team that hands out API keys to every engineer ends up paying the top-tier price for jobs a cheaper model could answer. Cloudflare said on 30 September its AI Gateway now has an Auto Router in public beta that picks a model per request based on a classifier running on Workers AI, and reports its own results as 30 percent cheaper than routing everything through OpenAI's GPT-6 Sol or Anthropic's Claude Opus 5.5.

The setup is small. Point a request at the model name cloudflare/auto through AI Gateway. The gateway reads a short slice of the recent conversation, runs it through a multi-head classifier that assigns probabilities across 14 task categories, rates the request on complexity, ambiguity, stakes and dependence on earlier context, then picks a model from the pool that survives the gateway's own filters (format support, credentials, spend limits, provider health).

The benchmark, from Cloudflare, against Cloudflare

Cloudflare says it uses Auto Router itself inside its OpenCode deployment and its Cloudflare OS agent harness. The published comparison is Cloudflare's own general-knowledge-work benchmark of 97 tasks, three samples per task per model, covering email, calendars, Slack, files, travel and finance.

Model Successful trials Success rate Total cost Cost per success
cloudflare/auto 252/291 86.6% $2.10 $0.0084
Anthropic Claude Opus 5.5 281/291 96.6% $5.91 $0.0210
OpenAI GPT-6 Sol 245/291 84.2% $2.64 $0.0108

Cloudflare reports 95 percent confidence intervals of +6.2 or -6.9 percentage points on Auto Router, +2.7 or -3.8 on Opus, and +6.5 or -6.9 on Sol, from 10,000 task-level bootstrap resamples. In dollar terms, Cloudflare's number is 80 percent of Sol and 35 percent of Opus. Opus still ranked highest on success at 96.6 percent, which the post is straight about: Auto Router does not beat the strongest model on a task, it declines to pay for the strongest model on tasks that a cheaper one clears.

What the tables leave out

Every figure in the table is Cloudflare benchmarking its own routing product against two competitors on its own benchmark, so read it as the vendor's number. The post does not publish per-category results, does not name every model in the routing pool, and does not describe how a task's difficulty was labelled to check the router's choice. It also acknowledges that a cheaper price per token can end up more expensive per outcome, because a weaker model uses more tokens per solve, and says the router optimises for predicted trajectory cost rather than dollars per million tokens.

The rules that let it work

The router runs on Workers AI on GPUs across Cloudflare's edge, so the classifier itself sits in the request path. Cloudflare says it filters out unhealthy upstream providers, brings them back after outages, and applies the gateway's existing controls: credentials, billing configuration, access policies and spend limits. Requests that need a specific model still get it: this is the routing default, not a lock.

For a team that has already put AI Gateway between its engineers and the providers (for spend visibility, per-user analytics, budget caps), Auto Router is one more setting on the same product. It moves the "does this task need Opus?" decision away from the individual writing the prompt, without blocking Opus for the work that needs it. Whether the 30 percent survives a team's own workload is a small experiment, run against your own traffic in front of the change, before trusting the Cloudflare benchmark.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX