Models & agents

Vercel says open-weight models now handle 56% of AI Gateway tokens, and Anthropic still takes 64% of the spend

September 18, 2026 at 10:00 AM PT

Isometric illustration of a retro-style computer monitor

Image: The New Stack

Why it mattersTwo teams reading the same headline can draw opposite conclusions about which lab their next production workload should target, and only one of them will be right.

Vercel published its September AI Gateway Production Index on Thursday. According to the report, open-weight models handled 56% of all tokens routed through the gateway in August, the first month they have taken the majority. Vercel says the same open-weight models accounted for just 14 cents of every estimated dollar spent through the gateway. The full breakdown was reported by Paul Sawers at The New Stack.

The token line

Vercel says open-weight models were 7% of gateway token volume in December 2025, 13% in April, 36% in July, and 56% in August. Every month between April and August rose. On August 22, Vercel CEO Guillermo Rauch said on LinkedIn that open-weight models had taken 62% of gateway traffic that day, a record. The gateway routes tens of trillions of tokens per month across applications running on Vercel.

The share of money going to those same models is a separate figure, and it moves the opposite way.

The spend line

Vercel says Anthropic took 64% of estimated dollar spend through the gateway in August. Its share has not fallen below 61% in any month since December 2025, and its models have held the top two positions by spend throughout that period, often the third as well. Open-weight models from labs like DeepSeek, Moonshot AI and Z.ai are cheaper to run, which is why a 56% share of tokens turns into 14% of dollars.

Vercel also says the average price per token across the gateway fell 23.2% in August, the third consecutive monthly decline. Among teams that processed more than 10 million tokens in both July and August, the median cost per token fell 7.6%.

Loyalty within a lab, disloyalty across labs

The most useful part of the report may be what it says about switching. Anthropic's Fable 5 fell from 13.2% of total gateway spend in July to 4.9% in August, while the cheaper Opus 5 climbed to 22.5%. Vercel says 90% of teams using Fable reduced their usage, and more of them moved to Opus 5 than to any other model. Vercel says Opus picked up almost twice as much usage as Fable lost, and attributes the shift to Opus handling similar workloads at about half the price.

The pattern reverses when the newer model does not carry the same capability at a lower price. Google's Gemini 3 Flash lost volume, and more than three-quarters of what it lost went to models from other labs, including OpenAI and Anthropic. Google's overall share of gateway token volume fell from 30% to 5% in one month, and Vercel says 22 of those 25 percentage points came from Gemini 3 Flash alone. The report authors write: "Lab loyalty doesn't follow brand, it follows model profile, and consistency wins."

Within five days of Z.ai launching GLM-5.3-Flash, Vercel says the new model was processing three times the daily volume of GLM-5.2.

The numbers come from one gateway's traffic, so they describe Vercel's customer base and not the whole market. Two things are worth taking from the report. Open-weight has passed 50% of gateway volume for the first time, and Anthropic still takes most of the money on it. A team quoting only the token share as a reason to migrate off a proprietary model, or a team quoting only the spend share as a reason to stay on one, is reading half the report.

Source

AI Gateway Production Index, September 2026 (Vercel). Reported by Paul Sawers in The New Stack.

Source: Vercel

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

Enclave says DeepSeek V4.1 Flash passed all 11 targets in its hacking benchmark for $4.65

The AI security firm Enclave says DeepSeek V4.1 Flash gained code execution on all 11 vulnerable targets in its hacking agent benchmark and left all four patched targets alone, at an accepted-run cost of $4.65 in API calls.

Source: Hacker NewsModels & agents

How Stale Is Your AI tracks 20 model families and shows that only 10 of them publish a training cutoff

A one-page site by Jock Mackinlay tracks the release date and the training cutoff for 20 current leading AI models side by side, and shows that only 10 of them publish a cutoff their lab actually documents.

Source: Hacker NewsModels & agents

Kuber Mehta argues Minecraft-in-one-prompt and the pelican-on-a-bicycle SVG are demo benchmarks that labs clearly prepare for by the next launch

A short essay by Kuber Mehta, at 73 Hacker News points on 6 September, argues that the viral one-prompt tests that follow every model launch are fixed public targets a lab has eight weeks to prepare for, and picks holdout evals as the alternative that still works.

Source: Hacker NewsModels & agents