Cloudflare's AI Gateway now tags each request by task and flags when a reasoning model is handling a job a smaller model could do
Cloudflare's AI Gateway User Insights now classifies every request by task and surfaces a "Model Overkill" view that flags when a reasoning model is answering a job a smaller model could handle, free for every AI Gateway user.

Image: Cloudflare
Why it mattersA team paying per token for a top-tier model is now one dashboard away from seeing how much of that spending is summarisation and formatting that a cheaper model would handle just as well.
Cloudflare has added a view to its AI Gateway that labels every request by the kind of work it is doing, and highlights the ones where a top-tier reasoning model is answering a question a smaller model could handle. The feature, announced on 30 September, is called User Insights, and Cloudflare says it is available to every AI Gateway user at no additional cost.
User Insights adds four categories to each request. Each one is classified by task, by model, by the number of conversation turns, and by the user who sent it. Cloudflare has started with five task labels: coding, research, writing, summarisation and data analysis. A conversation of ten turns about writing sits in a separate bucket from a one-shot summarisation.
Model Overkill and Potential Savings
The release ships two prebuilt views. "Model Overkill" surfaces requests where a high-capability reasoning model handled a job a smaller model could have done. "Potential Savings" puts a cost figure on each flagged request, so a team can see what the pattern is costing in the current month before deciding whether to change it. Cloudflare's own example is a user or an agent routing simple formatting requests to a reasoning model.
A closed beta of a feature called Auto Router goes further, routing future requests to the cheaper model automatically when the classifier decides a smaller model will do.
What the delay and the scope are
The classifier runs asynchronously, so new data appears in the dashboard about a day later rather than in real time. Cloudflare does not quote a cost figure for how much teams typically save with the new views, and the save estimates in Potential Savings are its own calculation against the gateway's logs. Cloudflare's own blog post is the source for both the feature list and the classifier categories.
A team that pays per token for a top-tier model has had no simple way to answer the question "how much of our spend is formatting and summarisation". The new view does not answer that for anyone outside Cloudflare's gateway, but inside it, a monthly report that shows the share of spend going into a reasoning model for jobs the gateway thinks a smaller model could handle gives a product manager a concrete number to argue over. Auto Router then takes the argument out of the loop, which is also where the gateway hides the model picker from the team.
Cloudflare also says the categories are a starting set and that it plans to add more. Teams running their own evals should check whether the classifier agrees with their own idea of what a request was doing before trusting the savings estimate.
Source
- Primary source: Cloudflare: Identify AI model overuse with User Insights, 30 September 2026
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


