MCP servers or command-line tools: what each adds to every message
An MCP server, a program that connects Claude Code to an outside service, adds its tool names and its instructions to every message. A command-line tool adds nothing until Claude runs it. That is the default behaviour in Anthropic's documentation, read on 4 October 2026: a feature named tool search holds back the full definition of each MCP tool until Claude needs it. When tool search is off, every definition loads at the start of the session and is sent with every request. Anthropic advises command-line tools such as `gh` and `aws` where they exist, because they add no tool listing. This page compares the two and shows how to check what your servers add.
Published October 4, 2026. Editorial.
Key takeaways
- With tool search on, which is the default, only MCP tool names and server instructions load at session start, by Anthropic's documentation read on 4 October 2026.
- In Anthropic's simulation of a session, the MCP tool names are 120 tokens, against 4,200 tokens for the system prompt.
- Setting `ENABLE_TOOL_SEARCH=false` loads every MCP tool definition at the start, and Claude Code does the same by default when `ANTHROPIC_BASE_URL` points to a server that is not Anthropic's own.
- Anthropic states that `gh`, `aws`, `gcloud` and `sentry-cli` use less context than MCP servers because they add no per-tool listing.
- Claude Code warns when one MCP tool result is over 10,000 tokens and limits a result to 25,000 tokens by default.
Claude Code is Anthropic's coding tool. You type a request, and an AI model reads files, runs commands and edits code for you. The work is counted in tokens. A token is a piece of text that the model processes, and every request you send is measured in tokens.
Claude Code can reach outside services, such as an issue tracker (a tool that lists a project's bugs and tasks) or a cloud account, in two ways. The first is an MCP server, a program that connects Claude Code to an outside tool or data source. The second is a command-line tool, a program that is run by typing its name in the terminal, which Claude can run by itself.
This page belongs to the guide on how to reduce Claude Code token usage. It compares what each of the two adds to every message. The short answer, from Anthropic's documentation read on 4 October 2026: with default settings an MCP server adds its tool names and its instructions to every request, and a command-line tool adds nothing until Claude runs it [1][2].
What an MCP server is, and what a tool definition is
MCP stands for Model Context Protocol. Anthropic's documentation, read on 4 October 2026, describes it as an open source standard for connecting AI to tools, and says MCP servers give Claude Code access to your tools, databases and other services [1]. The same page says when to connect one: when you find yourself copying data into the chat from another tool, such as an issue tracker or a monitoring dashboard [1].
Each server offers tools. A tool is one action that the model can ask Claude Code to perform, such as "search the issues". Before the model can call a tool, the request must contain the tool's definition. A tool definition is the text that tells the model what the tool does and what inputs it takes.
Definitions matter for cost because of their place in the request. The context window is the text the model can read in one request. The model keeps nothing between two requests, so Claude Code sends the full context again every time [4]. Tool definitions are in the system prompt layer, which is the first part of every request [4]. A definition that loads at the start of a session is therefore sent with every request until the session ends.
What an MCP server adds with tool search on
Claude Code has a feature named tool search. It holds back the full definition of each MCP tool until Claude needs it. Anthropic's documentation, read on 4 October 2026, says: "Only tool names and server instructions load at session start", and adds that Claude Code sets no fixed limit on the number of tools for each server [1]. Tool search is on by default [1].
So the fixed cost of a server, with default settings, has two parts:
- The tool names. In Anthropic's simulation of a session, the MCP tool names are 120 tokens, against 4,200 tokens for the system prompt [3]. Anthropic calls these representative counts, so your own figure depends on how many servers and tools you have.
- The server instructions. A server can send a text that tells Claude when to search for its tools. Claude Code cuts each server's instructions, and each tool description, at 2,048 characters by default [1].
When a task needs a tool, Claude searches for it and loads that one definition at that moment [3]. Tool search needs a model that supports it. The documentation lists Claude Sonnet 4.5, Claude Haiku 4.5, Claude Opus 4.5 and later models [1].
When every definition loads at the start
Tool search can be off, and you may not know it. When it is off, every MCP tool definition loads at the start of the session. This table lists the cases in Anthropic's documentation, read on 4 October 2026 [1].
| Situation | What loads at session start |
|---|---|
| Default settings on a supported model | Tool names and server instructions only |
ENABLE_TOOL_SEARCH=false |
Every MCP tool definition |
ENABLE_TOOL_SEARCH=auto |
Every definition while the definitions total less than 10% of the context window. Once they reach 10%, all are held back |
ENABLE_TOOL_SEARCH=auto:N |
The same, with your own percentage. auto:5 means 5% |
ANTHROPIC_BASE_URL points to a server that is not Anthropic's own |
Every definition, unless you set ENABLE_TOOL_SEARCH yourself |
CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS is set |
Every definition. Setting ENABLE_TOOL_SEARCH yourself cannot override this |
| Google Cloud's Agent Platform, on a model earlier than the Claude 4.5 generation | Every definition. ENABLE_TOOL_SEARCH=true does not override this |
| A Microsoft Foundry deployment hosted on Azure | Every definition. The deployment rejects tool search |
A server with alwaysLoad set to true |
Every tool of that server, whatever ENABLE_TOOL_SEARCH says |
ENABLE_TOOL_SEARCH is an environment variable, which is a named setting that a program reads when it starts. You can also set it in the env section of your settings.json file [1].
Three rows need a note.
The gateway row. A company can send Claude Code traffic through a gateway, a server of its own that receives each request from Claude Code and passes it on to the model. The address of that server goes in ANTHROPIC_BASE_URL. Anthropic's documentation says Claude Code turns tool search off in this case, because most such servers do not pass on the special blocks that tool search uses, named tool_reference blocks [1]. If your company uses a gateway, every MCP definition may be in every request you send. Setting ENABLE_TOOL_SEARCH=true turns tool search back on, and the documentation warns that requests then fail on a gateway that does not support those blocks [1]. The related guide has a page on gateways and cloud providers.
The auto row. In auto mode, definitions load at the start until they reach 10% of the context window. Anthropic's simulation uses a window of 200,000 tokens [3], and 10% of that is 20,000 tokens. That figure is our own arithmetic on Anthropic's numbers. It means that, with a window of that size, auto can place up to 20,000 tokens of definitions in every request before it starts to hold them back. The default setting holds them back from the first tool.
The alwaysLoad row. The documentation says to use this setting "for a small number of tools that Claude needs on every turn", because each tool loaded at the start uses context that would otherwise be available for your conversation [1].
What a command-line tool adds
Anthropic's cost page, read on 4 October 2026, gives this advice in a line that begins "Prefer CLI tools when available". CLI means command-line interface. The page names four tools, gh, aws, gcloud and sentry-cli, and says they use less context than MCP servers "because they don't add any per-tool listing". It adds: "Claude can run CLI commands directly." [2]
A command-line tool has no definition to load. Claude runs it the way it runs any other terminal command [2]. Because no listing is added for each program, ten installed programs add the same amount to the start of a session as zero.
The cost of a command-line tool comes later, as output. The result of every command enters the conversation and is then sent again with each later request. In Anthropic's simulation, one search across the code adds 600 tokens and one run of the tests adds 1,200 tokens, while the terminal shows a short status line [3]. A command that prints a long list costs more than a command that prints one line. Ask for only the result you need, and see the page on hooks that trim output for a way to filter output before Claude reads it. A hook is a command that Claude Code runs by itself at a fixed point, such as before every terminal command.
The limits on MCP output
An MCP tool result also enters the conversation. Anthropic's documentation, read on 4 October 2026, gives three limits [1]:
- Claude Code shows a warning when one MCP tool result is over 10,000 tokens. This warning level is fixed.
- The default maximum is 25,000 tokens. The
MAX_MCP_OUTPUT_TOKENSenvironment variable changes it, for exampleMAX_MCP_OUTPUT_TOKENS=50000. - A successful text result longer than 50,000 characters is saved to a file, whatever its token count, unless the server has raised the limit for that tool.
When a result is over the token limit, Claude Code saves it to a file and puts a message with the file path in the conversation, so Claude reads the file only when it needs the content [1].
Our advice is to leave MAX_MCP_OUTPUT_TOKENS at its default. A result of 25,000 tokens that enters the conversation is sent again with every later request, until you clear the conversation or Claude Code replaces it with a summary. Raising the limit allows larger results to do the same.
MCP server or command-line tool: the comparison
Every row comes from Anthropic's documentation, read on 4 October 2026 [1][2][3][4].
| MCP server | Command-line tool | |
|---|---|---|
| What loads at session start, default settings | Tool names and server instructions | Nothing. There is no per-tool listing |
| What loads at session start, tool search off | Every tool definition of every server | Nothing |
| When the rest loads | When Claude searches for a tool and uses it | No definition exists. Claude runs the command |
| Size in Anthropic's simulation | 120 tokens of tool names | No entry |
| What each use adds | The tool result, up to 25,000 tokens by default | The command output |
| How to switch it off | Toggle the server off in /mcp |
Nothing to switch off |
| Effect on the prompt cache of a change during a session | None with tool search on. With tool search off, adding a definition or removing one on purpose invalidates the cache | None |
| Anthropic's examples | Issue trackers, monitoring dashboards, databases | gh, aws, gcloud, sentry-cli |
How to see what your servers add, and reduce it
Start by measuring. Anthropic's cost page says: "Run /context to see what's consuming space." [2] The /context command shows the context window by category [3]. The page on how to read /usage and /context explains each screen.
Then open /mcp. The panel lists your servers and shows the number of tools next to each connected server [1]. Two facts from the documentation make this list worth reading [1]:
- You may have servers you did not add. If you log in to Claude Code with a claude.ai account, the MCP servers you added in claude.ai, which Anthropic calls connectors, are automatically available in Claude Code. Anthropic also provides some connectors itself. Setting
disableClaudeAiConnectorstotrueturns off the connectors that Claude Code fetches. - You can switch a server off and keep its configuration. Toggle it off in the
/mcppanel. Claude Code stops connecting to it and still lists it, marked as disabled. The choice is recorded for each project.
Anthropic's cost page gives the instruction in one line: run /mcp and disable any server you are not actively using [2]. On a Pro, Max, Team or Enterprise plan, the /usage screen also shows the share of recent usage attributed to each MCP server [2].
Subagents repeat the cost. A subagent is a second copy of Claude that works on one task in its own context window. In Anthropic's simulation, a subagent has the same MCP servers and skills as the main session, and that listing is 970 tokens in the subagent's own window [3]. A skill is a packaged set of instructions for one task.
Choose the moment for a change. The prompt cache is a store of request text that the service has already processed. When the start of a new request matches the store, that part is billed at a lower rate [4]. With tool search on, Claude Code keeps the tool list from the first request for the whole conversation, so a server that connects or disconnects during the session leaves the cache as it was. With definitions loaded at the start, adding a definition or removing one on purpose, for example by disabling a server in /mcp, invalidates the cache, and the next request is processed in full [4]. The related guide lists what invalidates the prompt cache. Choose your servers before you start a task.
Our position
Use the command-line tool when one exists for the service. This is Anthropic's own advice, and the reason is a count of what each one adds: a command-line tool adds nothing to a request until Claude runs it, and an MCP server adds its names and instructions to every request from the first one [2].
Keep MCP servers for services that have no command-line tool, and for the case the documentation describes, where you would otherwise copy data into the chat yourself [1]. For those servers, follow four rules:
- Leave
ENABLE_TOOL_SEARCHunset, so that every definition is held back. - If you work through a company gateway, run
/contextonce to check whether the definitions are loading at the start. - Turn off the servers a project does not use, in
/mcp, before you start work. - Leave
MAX_MCP_OUTPUT_TOKENSat 25,000.
Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau recommends checking the /mcp list at the start of each project, because a server that nobody uses still adds text to every request of every session. Reveneau is independent of Anthropic. Every figure on this page is Anthropic's own statement about its own product. The page on where Claude Code tokens go shows how the MCP listing compares with the rest of the context.
Common questions
What is an MCP server in Claude Code?
An MCP server in Claude Code is a program that connects Claude to an outside tool or data source, such as an issue tracker or a database. MCP stands for Model Context Protocol, which Anthropic's documentation, read on 4 October 2026, describes as an open source standard for connecting AI to tools. Each server offers tools, and the model must receive a tool's definition before it can call that tool.
What is MCP tool search?
MCP tool search is the Claude Code feature that holds back the full definition of each MCP tool until Claude needs it. Anthropic's documentation, read on 4 October 2026, says only tool names and server instructions load at session start, and Claude loads a specific definition when a task needs it. Tool search requires Claude Sonnet 4.5, Claude Haiku 4.5, Claude Opus 4.5 or a later model.
Is MCP tool search on by default?
Yes. MCP tool search is on by default in Claude Code, by Anthropic's documentation read on 4 October 2026. With nothing set, Claude Code still loads every definition at the start in three documented cases: when `ANTHROPIC_BASE_URL` points to a server that is not Anthropic's own, on Google Cloud's Agent Platform models earlier than the Claude 4.5 generation, and on a Microsoft Foundry deployment hosted on Azure.
When does Claude Code load every MCP tool definition at the start?
Claude Code loads every MCP tool definition at the start of the session when tool search is off. Among the causes in Anthropic's documentation, read on 4 October 2026, are `ENABLE_TOOL_SEARCH=false`, the `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS` variable, a gateway address that is not Anthropic's own, and a Microsoft Foundry deployment hosted on Azure. A server marked `alwaysLoad: true` also loads all of its own tools at the start.
How do I see how much context my MCP servers use?
Run `/context` to see how much of the context window your MCP servers use. Anthropic's cost documentation, read on 4 October 2026, names `/context` as the command that shows what is using space. The `/mcp` panel shows the number of tools next to each connected server. On a Pro, Max, Team or Enterprise plan, `/usage` also shows the share of recent usage attributed to each MCP server.
How do I turn off an MCP server without deleting it?
Open the `/mcp` panel and toggle the MCP server off. Anthropic's documentation, read on 4 October 2026, says this stops Claude Code from connecting to the server and keeps its configuration, and the server stays in the `/mcp` list marked as disabled. Claude Code records the choice per project. Anthropic's cost page advises disabling any server you are not actively using.
Why does Anthropic recommend command-line tools over MCP servers?
Anthropic recommends command-line tools where they exist because they add no per-tool listing to the context. Its cost documentation, read on 4 October 2026, names `gh`, `aws`, `gcloud` and `sentry-cli`, says they use less context than MCP servers, and says Claude can run command-line commands directly. An MCP server adds its tool names and instructions to every request, even with tool search on.
Is there a limit on the size of MCP tool output?
Yes. Claude Code limits one MCP tool result to 25,000 tokens by default and shows a warning when a result is over 10,000 tokens, by Anthropic's documentation read on 4 October 2026. A result over the limit is saved to a file, and the conversation receives a message with the file path. The `MAX_MCP_OUTPUT_TOKENS` variable changes the limit, and the warning level is fixed.
What does ENABLE_TOOL_SEARCH=auto do?
`ENABLE_TOOL_SEARCH=auto` loads MCP tool definitions at the start while they total less than 10% of the context window, and holds all of them back once they reach 10%. Anthropic's documentation, read on 4 October 2026, calls this threshold mode. The form `auto:5` sets the threshold to 5%. With the variable left unset, all MCP tools are held back, which is the smaller start-up cost.
Does turning off an MCP server during a session affect the prompt cache?
Turning off an MCP server during a session affects the prompt cache only when tool definitions were loaded at the start. Anthropic's documentation, read on 4 October 2026, says that with tool search on, Claude Code keeps the tool list from the first request for the whole conversation. With tool search off, disabling a server in `/mcp` removes its definitions, which invalidates the cache, so the next request is processed in full.
Why does /mcp list servers I never added?
The `/mcp` list can include MCP servers from your claude.ai account. Anthropic's documentation, read on 4 October 2026, says servers added in claude.ai, called connectors, are automatically available in Claude Code when you log in with a claude.ai account, and Anthropic provides some connectors itself. To turn off the connectors Claude Code fetches, set `disableClaudeAiConnectors` to `true`, or toggle one off in `/mcp` for the current project.
Should I set alwaysLoad on an MCP server?
Set `alwaysLoad` on an MCP server only when Claude needs its tools on every turn. Anthropic's documentation, read on 4 October 2026, says every tool from such a server loads into the context at session start, whatever `ENABLE_TOOL_SEARCH` is set to, and that each tool loaded this way uses context that would otherwise be available for your conversation. It advises the setting for a small number of tools.
References
- Anthropic, Connect Claude Code to tools via MCP (code.claude.com), read 4 October 2026
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Explore the context window (code.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
More in Setup
How long should CLAUDE.md be? What to keep and what to move into skills
A CLAUDE.md file should be under 200 lines. That is the target in Anthropic's documentation for Claude Code, read on 4 October 2026, and the reason is cost: CLAUDE.md loads at the start of every session and is then sent with every request, so each line is counted on every request, including requests that have no use for it. Keep the facts that every session needs, such as build commands and conventions. Move step-by-step procedures into skills, which load when they are used, and move instructions for one folder into path rules, which load when Claude opens a matching file. This page shows where to put each type of instruction.
Hooks that trim test and log output before Claude reads it
A hook can remove the passing lines from test output, and the lines without errors from a log, before Claude reads them. A hook is a command that Claude Code runs by itself at a fixed point, and a hook on the PreToolUse event can rewrite a terminal command before it runs. Anthropic's documentation, read on 4 October 2026, gives a working example that keeps only the failing lines of a test run. This page reproduces that example, shows how to check it with `/hooks` and a debug log, and lists what the filter can hide, such as a failure that is reported with an unexpected word. It also covers code intelligence plugins and subagents.