Limits

Rate limits per user by team size: how many tokens per minute to ask for

Anthropic's documentation, read on 4 October 2026, recommends a rate limit per Claude Code user that falls as the team grows: 200,000 to 300,000 tokens per minute for a team of 1 to 5 users, 15,000 to 20,000 for a team of 100 to 500, and 10,000 to 15,000 above 500 users. A rate limit is the most work an organisation may send to the model in one minute, measured in tokens (pieces of text) and in requests. Anthropic applies these limits at the organisation level, so one developer may use more than their share while others are idle. This page reproduces Anthropic's table, adds the organisation-wide total for each team size, and says which figure to ask for.

Published October 4, 2026. Editorial.

Key takeaways

  • Anthropic's documentation, read on 4 October 2026, recommends 200,000 to 300,000 tokens per minute per user for a Claude Code team of 1 to 5 users, falling to 10,000 to 15,000 per user above 500 users.
  • Anthropic's own worked example is a team of 200 users at 20,000 tokens per minute each, which is 4 million tokens per minute for the organisation.
  • The per-user figure falls as the team grows because, in Anthropic's words, fewer users tend to use Claude Code concurrently in larger organisations.
  • The limits apply at the organisation level, so an individual can temporarily use more than their calculated share when others are idle, by Anthropic's documentation.
  • Anthropic says a workspace rate limit can be set on the Claude Code workspace in the Claude Console to cap Claude Code's share of the organisation's API limits and protect other production workloads.

A rate limit is the most work an organisation may send to the model in one minute. It is separate from a spend limit, which is a cap in dollars. The rate limit decides how many developers can work at the same moment before some of them have to wait. Anthropic's documentation on managing costs, read on 4 October 2026, gives a table of recommended rate limits per Claude Code user for each team size, and this page reproduces it [1].

Two units appear in that table. A token is a piece of text that the model reads or writes; Anthropic's pricing page, read on 4 October 2026, puts one token at 4 characters or 0.75 of an English word [3]. Tokens per minute, which Anthropic abbreviates to TPM, is the amount of text an organisation may send and receive in one minute. Requests per minute, abbreviated to RPM, is the number of separate calls to the model in one minute. Each time Claude Code sends the conversation to the model and gets an answer back, that is one request, and one request can contain many thousands of tokens.

This page is part of the guide to Claude Code costs for teams. It covers the Claude Console setup, where an organisation holds an account with Anthropic and is billed per token, because that is where Anthropic's table applies. On a Claude for Teams or Enterprise plan the control is different: each member has a seat, a paid place on the plan, and each seat comes with an allowance, an amount of usage included in the seat's price that resets on a rolling five-hour window and a weekly window [1]. The page on subscription seats or pay per token explains the difference between the two setups.

Anthropic's recommendation table, with the organisation-wide total

Anthropic's table gives a tokens-per-minute figure and a requests-per-minute figure per user for six team sizes. The first three columns below reproduce it exactly as it appears in the documentation read on 4 October 2026 [1]. The fourth column is this page's own arithmetic: the number of users at the top of each row multiplied by the higher tokens-per-minute figure in that row. Anthropic's documentation gives only the 200-user example in the next section, so every figure in the fourth column is Reveneau's multiplication, shown so a manager can see the organisation-wide number that a per-user figure implies.

Team size (Anthropic) TPM per user (Anthropic) RPM per user (Anthropic) Organisation total at the top of the row (this page's arithmetic)
1-5 users 200k-300k 5-7 5 users at 300,000 is 1,500,000 tokens per minute
5-20 users 100k-150k 2.5-3.5 20 users at 150,000 is 3,000,000 tokens per minute
20-50 users 50k-75k 1.25-1.75 50 users at 75,000 is 3,750,000 tokens per minute
50-100 users 25k-35k 0.62-0.87 100 users at 35,000 is 3,500,000 tokens per minute
100-500 users 15k-20k 0.37-0.47 500 users at 20,000 is 10,000,000 tokens per minute
500+ users 10k-15k 0.25-0.35 No upper end; 500 users at 15,000 is 7,500,000 tokens per minute

Three things to read in that table. First, Anthropic's rows share their boundary numbers: a team of exactly 5, 20, 50, 100 or 500 users appears in two rows. The documentation gives no rule for the boundary, so this page's totals use the higher figure at each boundary, and a team at a boundary should expect the two readings to differ. Second, the requests-per-minute figures fall below one per user from the 50 to 100 row onwards. A figure of 0.62 requests per minute per user only makes sense as a share of one organisation-wide number, and the next sections explain why Anthropic sizes the limit that way. Third, the last row has no upper end, so the figure shown for it is the total at 500 users; a team of 1,000 users at 15,000 each would be 15,000,000 tokens per minute, by the same arithmetic.

Anthropic's 200-user example

Anthropic's documentation gives one worked example under the table. For a team of 200 users, it says, you might request 20,000 tokens per minute for each user, or 4 million total tokens per minute, and it shows the multiplication: 200 times 20,000 is 4 million [1]. Two details of that example are worth copying. Anthropic took the top of the 100 to 500 row, 20,000, and it multiplied by the whole headcount, every user who has access, with no reduction for people who are away. The request to Anthropic is for the organisation-wide figure, 4 million, and the per-user figure is only the way to arrive at it.

Why the per-user figure falls as the team grows

A manager reading the table will ask why a developer in a team of 300 is given a fifteenth of what a developer in a team of 3 is given, comparing the top figure of each row. Anthropic's answer, in its documentation read on 4 October 2026, is that the tokens-per-minute figure per user decreases as the team grows because fewer users tend to use Claude Code concurrently in larger organisations [1]. In a team of 3, all three may be working at the same minute. In a team of 300, at any one minute some are in meetings, some are reading code, and some are in a different time zone. The organisation-wide limit only has to cover the people sending requests at the same moment, so the figure per head can fall as the head count rises.

That reasoning applies only while the pattern applies. Anthropic adds a note that if you anticipate scenarios with unusually high concurrent usage, such as live training sessions with large groups, you may need higher tokens-per-minute allocations per user [1]. A training day, a hackathon, or a day when a whole department is told to try the tool at once does not match the assumption behind the table, and Reveneau recommends asking for the higher figure before such a day instead of after it.

The limit belongs to the organisation, and one developer can use more than their share

Anthropic's documentation states that these rate limits apply at the organisation level, and that this means individual users can temporarily consume more than their calculated share when others are not actively using the service [1]. For a manager this has two consequences. The per-user figure in the table is a planning number for sizing the request to Anthropic; the limit itself is enforced on the organisation as a whole. And when every developer does work at once, the organisation-wide limit is what they share, so a team that sized its request from a quiet week will find developers waiting on a busy one.

The place where Claude Code's share can be controlled is a workspace. In the Claude Console, a workspace is a named section of the organisation's account with its own usage figures and its own limits. Anthropic's documentation says that when Claude Code first authenticates with a Claude Console account, a workspace called Claude Code is created automatically, that it provides centralised cost tracking for all Claude Code usage in the organisation, and that no API keys can be created for it [1]. For organisations with custom rate limits, Claude Code traffic in this workspace counts toward the organisation's overall API rate limits, and Anthropic says a workspace rate limit can be set on this workspace's Limits page in the Claude Console to cap Claude Code's share and protect other production workloads [1].

That last sentence matters to any company whose Claude Console account also serves a product. Without a workspace rate limit, a busy afternoon of Claude Code use and the company's own application draw on one organisation-wide limit. With one, Claude Code is held to the share the administrator chooses. Anthropic's pricing page, read on 4 October 2026, says rate limits vary by usage tier, which it names as Start, Build and Scale, and that limits beyond the Scale tier and custom rate limits are arranged with Anthropic's sales team [3]. So the request for a higher organisation-wide figure and the cap on Claude Code's share of it are two separate steps, one with Anthropic and one on the Limits page.

What a developer sees when a rate limit is reached

A developer who reaches a rate limit sees an error. Anthropic's error reference, read on 4 October 2026, gives the text as Request rejected (429), with a trailing sentence that names where to check service health, and says it means the rate limit configured for the API key, Amazon Bedrock project or Google Cloud project has been reached [2]. Anthropic's steps are to run /status to confirm that the active credential is the one expected, because a stray API key in the environment can route requests through a low-tier key instead of the subscription; to check the provider console for the active limits and request a higher tier if needed; and to reduce concurrency, by lowering the CLAUDE_CODE_MAX_TOOL_USE_CONCURRENCY environment variable, by running fewer parallel subagents, or by switching to a smaller model for high-volume scripted runs [2]. An environment variable is a named value the operating system passes to a program when it starts; a subagent is a second copy of Claude that does one task in its own separate conversation and returns a summary.

A second message needs its own explanation. Anthropic's error reference describes Server is temporarily limiting requests as a short-lived throttle on Anthropic's side, unrelated to the plan's quota, and the message text includes the words not your usage limit. As of Claude Code version 2.1.199, Claude Code retries it automatically with backoff before showing it, whichever way the session authenticates [2]. The messages that begin with the words You've hit your are a third kind: they report seat allowances and spend limits, and the page on what each limit message means covers them.

Which figure to ask for

Reveneau's recommendation, from Anthropic's figures, is to take the row for your head count, use the higher tokens-per-minute figure in it, and multiply by every developer who will have access, which is the method Anthropic's own 200-user example follows [1]. For a team at a boundary between two rows, take the higher reading. Then, if the organisation's account also serves production software, set a workspace rate limit on the Claude Code workspace so that developers and the product do not draw on one limit [1]. Revisit the figure before any day on which the whole team will use the tool at once, because that is the case Anthropic's table does not cover.

A rate limit sets how fast the team can work. Spending is capped separately, by the spend controls on the page about spend limits for an organisation, a group and a person, and the overall figures a budget rests on are on the page about what Claude Code costs per developer. A developer who wants to send fewer tokens per request, and so fit more work inside any limit, will find the habits in the sister guide on how to reduce Claude Code token usage.

Reveneau is an AI software development consultancy. All of its code is written by AI, and every change must pass an eval suite, a set of automated tests written from the specification, before release, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic, and every figure on this page is either Anthropic's own statement about its own product or this page's labelled arithmetic on those statements. If you want help sizing a rate limit request for your team, contact Reveneau.

Common questions

What is a tokens-per-minute limit for Claude Code?

A tokens-per-minute limit is the largest amount of text, measured in tokens, that an organisation may send to and receive from the model in one minute through its Claude Console account. A token is a piece of text that Anthropic's pricing page, read on 4 October 2026, puts at 4 characters or 0.75 of an English word. Anthropic's recommendation table gives this figure per user, from 200,000 to 300,000 for a team of 1 to 5 users.

What is a requests-per-minute limit for Claude Code?

A requests-per-minute limit is the largest number of separate calls to the model that an organisation may make in one minute. Each time Claude Code sends the conversation to the model, that is one request. Anthropic's documentation, read on 4 October 2026, recommends 5 to 7 requests per minute per user for a team of 1 to 5 users and 0.25 to 0.35 per user above 500 users, as shares of one organisation-wide limit.

How many tokens per minute does Anthropic recommend per user for a team of 200?

For a team of 200 users, Anthropic's documentation, read on 4 October 2026, places the team in its 100 to 500 user row, which recommends 15,000 to 20,000 tokens per minute per user and 0.37 to 0.47 requests per minute per user. Anthropic's own worked example takes the top of that row: 200 users at 20,000 tokens per minute each, which it states as 4 million tokens per minute in total.

Why does Anthropic's recommended tokens per minute per user fall as the team grows?

Anthropic's recommended tokens per minute per user falls as the team grows because, in the words of its documentation read on 4 October 2026, fewer users tend to use Claude Code concurrently in larger organisations. The limit is one organisation-wide number, so it only has to cover the people working at the same moment. A team of 1 to 5 users gets 200,000 to 300,000 per user; a team above 500 gets 10,000 to 15,000 per user.

Do Claude Code rate limits apply to each developer separately?

No. Anthropic's documentation, read on 4 October 2026, says the recommended rate limits apply at the organisation level. The per-user figures in Anthropic's table are a way to size the organisation-wide total: multiply the figure by the number of users. That is why one developer can temporarily consume more than their calculated share when other developers are not actively using the service.

What total tokens per minute should a team of 50 developers request for Claude Code?

A team of 50 developers falls at the boundary of two rows in Anthropic's table, read on 4 October 2026: 20 to 50 users at 50,000 to 75,000 tokens per minute per user, and 50 to 100 users at 25,000 to 35,000. This page's own arithmetic on those figures gives 3.75 million tokens per minute at 50 users and 75,000 each, or 1.75 million at 50 users and 35,000 each. Reveneau recommends asking for the higher figure.

Do live training sessions need a higher Claude Code rate limit?

Yes, by Anthropic's documentation read on 4 October 2026, which says that scenarios with unusually high concurrent usage, such as live training sessions with large groups, may need higher tokens-per-minute allocations per user. The reason follows from Anthropic's table: the per-user figures assume that only some of the team works at the same moment, and a training session puts everyone on the tool at once.

What is a workspace rate limit in the Claude Console?

A workspace rate limit is a cap on how much of the organisation's overall rate limit one workspace, a named section of a Claude Console account, may use. Anthropic's documentation, read on 4 October 2026, says that for organisations with custom rate limits, Claude Code traffic in the automatically created Claude Code workspace counts toward the organisation's overall API rate limits, and that a workspace rate limit can be set on that workspace's Limits page to cap Claude Code's share.

Does Claude Code traffic count against our other API rate limits?

Yes, for organisations with custom rate limits on the Claude Console. Anthropic's documentation, read on 4 October 2026, says Claude Code traffic in the Claude Code workspace counts toward the organisation's overall API rate limits. If the same account also serves a production application, Claude Code and that application share one limit. Anthropic's stated remedy is a workspace rate limit on the Claude Code workspace, set on its Limits page, to protect other production workloads.

What does Request rejected (429) mean in Claude Code?

Request rejected (429) means Claude Code has reached the rate limit configured for the API key, Amazon Bedrock project or Google Cloud project it is using, by Anthropic's error reference read on 4 October 2026. Anthropic's steps are to run /status to confirm the active credential is the expected one, check the provider console for the active limits and request a higher tier if needed, and reduce concurrency, for example by running fewer parallel subagents.

Is Server is temporarily limiting requests the same as a rate limit?

No. Anthropic's error reference, read on 4 October 2026, describes Server is temporarily limiting requests as a short-lived throttle on Anthropic's side that is unrelated to the plan's quota, and the message text includes the words not your usage limit. As of Claude Code version 2.1.199, Claude Code retries it automatically with backoff before showing it, whichever way the session authenticates. Anthropic's advice is to wait briefly and check status.claude.com if it persists.

Does Anthropic's rate limit table apply on a Teams or Enterprise plan?

Anthropic's rate limit table appears in the Claude Console section of its cost documentation, read on 4 October 2026, and is introduced as advice for setting up Claude Code for teams on that setup. On a Claude for Teams or Enterprise plan, the documentation describes a different control: each member's usage draws from a seat allowance that resets on a rolling five-hour window and a weekly window. The documentation gives no tokens-per-minute table for those plans.

More in Limits