Infrastructure

OpenRouter turns on US in-region routing, and DeepSeek, Kimi and GLM traffic can now stay entirely inside the United States

September 14, 2026 at 7:35 AM PT

OpenRouter blog card reading In-Region Routing: Keep your data in the US or EU

Image: OpenRouter

Why it mattersA team that wanted to use DeepSeek V4 Pro, Kimi K3 or GLM 5.2 through a router but could not clear a US data residency review now has an endpoint that answers the review with a yes.

OpenRouter has made US in-region routing generally available for its Business and Enterprise customers, alongside the EU version it shipped in October 2025. Requests sent to us.openrouter.ai are decrypted inside the United States, routed only to provider endpoints running in the United States, and rejected with a 404 if no compliant provider can serve the requested model. The company published the release note on 9 September, and The New Stack picked it up today, on 14 September.

Why the endpoint exists

Cailee Moberg, on OpenRouter's product team, writes on the company blog that open-weight models now account for 60 percent of the tokens US-originating requests send through OpenRouter, up from 26 percent a year earlier. Most of that volume is Chinese: DeepSeek V4 Pro, Kimi K3, and GLM 5.2 are the three the company names.

"Models from Chinese labs are still most of the [open-weight model] volume, and procurement approval for those models can be difficult," Moberg writes. OpenRouter's argument is that when a US or EU provider hosts a Chinese-weight model, "requests go to that provider and the lab is not involved," so a team that could not get past a data residency review before now can.

The company also cites a 2026 Deloitte survey figure, that 77 percent of enterprises "now factor country of origin into their vendor selection," as the market backdrop.

What changes at the routing layer

Switching is one URL swap: https://openrouter.ai/api/v1 becomes https://us.openrouter.ai/api/v1, and the API key, request body, and model IDs stay the same. OpenRouter decrypts the request inside the region, filters the pool of providers down to those it has approved as running there, and hides everything else for that request. Guardrails at the workspace, team, or API-key level can enforce the choice, so a wrong hostname gets rejected with an error instead of a silent downgrade.

Moberg distinguishes this from what she calls inference-only regional routing, where a gateway pins the provider's inference to a region but the request is still decrypted somewhere else first. On OpenRouter's US and EU endpoints, the prompt never exists in plaintext outside the region, and features that would send prompt data outside the region, such as some web-search tools, are disabled on the regional endpoint.

What is available today

DeepSeek V4 Pro, Kimi K3, and GLM 5.2 route through the US endpoint because Baseten, Fireworks, and Azure host them from US data centres. GLM 5.2 also has an EU endpoint through Mistral. The US list also includes GPT-5.6, Claude Opus 5, Gemini 3.6 Flash, Grok 4.6, NVIDIA's Nemotron 3 Ultra, Thinking Machines' Inkling, gpt-oss-120b, and Qwen3 Coder. The live catalogue is at openrouter.ai/models?region=us, and the same page swapped to region=eu shows the European list.

Where this fits

OpenRouter sits between an application and hundreds of hosted models, and picks a provider for each request. Stripe agreed to acquire the company for a reported $8 billion, and Cursor, Ramp, and Meta are all building their own routers, according to The New Stack. What OpenRouter is doing here is turning that routing layer, the thing that already decides which provider serves each call, into a compliance surface: a US customer that wants the price and quality of a Chinese open-weight model can get it without also getting a data residency question they cannot answer.

Source

Primary source: In-Region Routing: Keep your data in the US or EU on the OpenRouter blog, by Cailee Moberg, 9 September 2026. Reporting: Paul Sawers at The New Stack, 14 September 2026.

Source: OpenRouter

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

Engineer Jyn says GLM 5.3-flash puts capable offensive AI on a $9,500 Mac, and security teams have about a year to prepare

An engineer writing at jyn.dev argues that GLM 5.3-flash and its refusal-stripped variants have made capable offensive AI cheap enough to run at home, and that industry has about a year before this becomes routine.

Source: Hacker NewsInfrastructure

Nvidia is reportedly buying Hugging Face for $12.9 billion, and neutrality is most of what it is buying

The Information reported that Nvidia agreed to buy Hugging Face for $12.9 billion. At roughly $150 million in annualised revenue, that is a multiple of about 86, so the price is for position rather than earnings.

Source: PressInfrastructure

GLM-5.3 went open weight and dropped MIT: hosts above $10 billion in revenue now need a security review

Z.ai put GLM-5.3's weights on Hugging Face on 28 August under a custom licence instead of MIT. Individuals are unaffected. Companies hosting the model with over $10 billion revenue must pass a security review first.

Source: PressModels & agents