When a subagent saves tokens in Claude Code, and when it adds them
A subagent saves tokens in your main conversation when the task produces output you will never read again: a test run, a log file, a documentation fetch. A subagent is a second copy of Claude with its own conversation, its own system prompt and its own tools, and only its summary comes back. Anthropic's documentation, read on 4 October 2026, says the subagent's own requests still count toward the same usage limits as the main conversation. So the saving is in the main conversation, and the subagent's start-up and work are paid in full. This page explains when that trade is worth it, what a subagent loads before it works, why its first request reads none of the parent's prompt cache, and how to prompt it.
Published October 4, 2026. Editorial.
Key takeaways
- A subagent runs in its own context window with its own system prompt and tools, and only its final text returns to the main conversation, by Anthropic's documentation read on 4 October 2026.
- In Anthropic's own simulation, a subagent reads 6,100 tokens of files and returns a 420-token result, so the main conversation grows by 420 tokens.
- The subagent's requests count toward the same usage limits as the main conversation, and the attribution section of /usage shows the subagent share on a Pro, Max, Team or Enterprise plan.
- A subagent's first request reads none of the parent's prompt cache, because its prefix differs; a fork inherits the parent's system prompt, tools and history and reads the parent's cache.
- By default 20 subagents can run at once, and a subagent can start its own subagents up to three layers below the main conversation, both figures from Anthropic's documentation.
Claude Code is Anthropic's coding tool: you type a request in a terminal, and an AI model reads files, runs commands and edits code for you. Everything the model reads and writes is counted in tokens, which are pieces of text. All of that text goes into the context window, which is the full set of text the model reads on each request. A subagent is a second copy of Claude that takes one task into a separate context window of its own.
This page is part of the guide to the cost of agents and unattended runs in Claude Code. It answers one question: when does a subagent save tokens, and when does it add them? The short answer, from Anthropic's documentation read on 4 October 2026, is that a subagent saves tokens in the main conversation and spends tokens of its own, and the two have to be compared.
What a subagent is, and what comes back
Anthropic's documentation describes a subagent in one paragraph. Each subagent "runs in its own context window with a custom system prompt, specific tool access, and independent permissions. It also sends its own requests, which count toward the same usage limits as your main conversation." [2] The system prompt is the set of instructions the model reads first, and a subagent's system prompt is its own, shorter than the main session's [5].
The reason to use one is stated as plainly: Anthropic says to use one when a side task would fill your main conversation with search results, logs or file contents you will never read again, because the subagent does that work in its own context and returns only the summary [2].
Only the summary returns. Anthropic's context window page, which simulates a session step by step, says that "Only the subagent's final text response comes back to your context, plus a small metadata trailer with token counts and duration." [5] The file reads, the command output and the subagent's own reasoning stay in the subagent's transcript, which Claude Code stores as a separate file [2].
Where the saving is: the main conversation
The main conversation is sent again to the model with every request, and each time Claude uses a tool it sends another request that holds the results [1]. Anything that enters the main conversation is therefore paid for again on every later turn, at the cached token rate once the prompt cache holds it. The prompt cache is a store, kept by the service that runs the model, of request text it has already processed, so that unchanged text is billed at a lower rate [3].
That is why a subagent saves tokens. Anthropic's simulation gives the figures. The subagent in it reads three sets of files of 2,200, 800 and 3,100 tokens, which is 6,100 tokens in total, and returns a 420-token result [5]. The main conversation grows by 420 tokens instead of 6,100, and every later request in the main conversation includes 420 tokens in place of 6,100. Our own arithmetic on Anthropic's figures: the difference is 5,680 tokens on each later request.
Anthropic's cost page names the tasks this suits: "Running tests, fetching documentation, or processing log files can consume significant context. Delegate these to subagents so the verbose output stays in the subagent's context while only a summary returns to your main conversation." [1] The subagent page gives the prompt to use: "Use a subagent to run the test suite and report only the failing tests with their error messages" [2].
What the subagent pays
The subagent's work is paid in full. Anthropic's cost page says it directly: "The subagent's own requests still draw on your usage." [1] On a Pro, Max, Team or Enterprise plan, the attribution section of /usage shows recent usage attributed to skills, subagents, plugins and MCP servers, each as a percentage of the total, computed from the session history on this machine [1]. The page on how to read /usage and /context covers that screen.
A subagent starts with text of its own before it reads a single file. Anthropic lists what a non-fork subagent's initial context contains [2]:
- Its system prompt, which is the agent's own prompt plus environment details, in place of the Claude Code system prompt.
- The task message, which is the delegation prompt Claude writes when it hands off the work.
- CLAUDE.md files, the instruction files you write for Claude, at every level the main conversation loads. The built-in Explore and Plan agents skip these. A custom subagent skips the user, project and local files when its definition sets
omitClaudeMd: true, which requires Claude Code v2.1.271 or later. - Git status, a snapshot of the repository, which Explore and Plan also skip.
- Preloaded skills, in full, when the agent's
skillsfield names any. A skill is a packaged set of instructions for one task.
In the simulation, this start-up text is 900 tokens of system prompt, 1,800 tokens of project CLAUDE.md, 970 tokens of MCP tools and skills, and a 120-token task prompt [5]. An MCP server is a program that connects Claude Code to an outside tool or data source. Added together, our arithmetic on Anthropic's figures gives 3,790 tokens before the first file read. The main conversation's auto memory, the notes Claude writes for itself, is left out [2].
Two more costs apply when there are many subagents. Anthropic warns that "Running many subagents that each return detailed results can consume significant context, and each subagent spends tokens of its own while it runs." [2] And a subagent can start subagents of its own, by default up to three layers below the main conversation; at the depth limit Claude Code withholds the Agent tool, which is the tool Claude uses to start a subagent, and CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH changes the limit, with 1 turning nesting off [2]. Each nested subagent sends its own requests.
There is also a limit on how many run at once. By default, when 20 subagents are running, starting another fails with Concurrent subagent limit reached; the limit is set with CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS, requires Claude Code v2.1.217 or later, and sessions with ultracode on, a setting under which Claude runs large tasks as workflows, are exempt from it. There is no limit on the total number over a session [2].
The first request reads none of the parent's cache
A subagent's prompt cache starts empty. Anthropic's prompt caching page says: "A subagent starts its own conversation with its own system prompt and tool set, separate from the parent's. Its first request doesn't read the parent's cache, because the two prefixes differ", and it then builds a cache of its own across its turns [3]. The prefix is the start of the request, which the cache matches exactly. So the subagent's first request reads nothing from any cache and is processed in full, and later requests in the same subagent read its own cache.
The parent is unaffected: "From the parent's side, the subagent's call and result append to the conversation, leaving the parent's prefix intact." [3] Subagent requests also fall outside the main conversation's cache lifetime bucket, so they get a five-minute lifetime even on a subscription unless you choose a longer one; the page on the five-minute and one-hour cache lifetimes covers that setting.
A fork is the exception. A fork is a subagent that inherits the entire conversation so far, with the same system prompt, tools, model and message history as the main session [2]. "Because a fork's system prompt and tool definitions are identical to the parent, its first request reuses the parent's prompt cache." Anthropic adds that this makes forking cheaper than starting a new subagent for tasks that need the same context [2]. You can start one yourself with /subtask followed by a task, which requires Claude Code v2.1.212 or later [2].
A skill can also run in a subagent. Add context: fork to a skill's frontmatter, the settings block at the top of its file, and Claude Code starts a new subagent of the type in the skill's agent field, with the skill content as its prompt. Anthropic notes that despite the name, this subagent does not see your conversation history, so the skill's instructions have to stand on their own; when the task depends on the history, fork the conversation instead [4]. A forked skill using agent: Explore sees only the skill content and the agent's own system prompt, because Explore skips CLAUDE.md and git status [4].
Resuming helps too. Each subagent invocation normally creates a new instance. Claude can instead resume a finished subagent by name or agent ID with the SendMessage tool, and the resumed run keeps its full history and can keep reading the prompt cache that the original run built [2]. Explore and Plan return no agent ID and cannot be resumed [2].
Write a focused spawn prompt
Spawn is Anthropic's verb for starting an agent, and the spawn prompt is the message that starts it. The task message Claude writes is part of the subagent's initial context, so everything in it is read on the subagent's first request and on every request after. Anthropic's cost page states the rule for agent team teammates: "Keep spawn prompts focused. Teammates load CLAUDE.md, MCP servers, and skills automatically, but everything in the spawn prompt adds to their context from the start." [1] A subagent's task message enters its context the same way, by the start-up list above, so Reveneau recommends the same practice for subagents.
Two things belong in the prompt. First, the task and the shape of the answer you want, as in Anthropic's example that asks for "only the failing tests with their error messages" [2]. Second, any rule from your CLAUDE.md that the subagent itself must follow. Anthropic says most rules do not need to reach the subagent, because the main conversation still has your full CLAUDE.md when it reads the result; a rule that must reach it, such as "ignore the vendor/ directory", should be restated in the delegation prompt [2].
One thing belongs outside the prompt: the subagent's description field. Claude reads every subagent's description to decide when to delegate, and those descriptions take up context in the main conversation. When the combined descriptions of your own subagents, excluding the built-in ones, exceed 15,000 tokens, Claude Code shows a warning at startup. Anthropic's advice is to keep descriptions short and move detail into the subagent's system prompt, which loads only when that subagent runs [2].
Which tasks to delegate
Anthropic's documentation lists when to use the main conversation and when to use a subagent [2]. The table restates that list with the reason for each row.
| Task | Main conversation or subagent | Why, by Anthropic's documentation |
|---|---|---|
| Run the test suite and report failures | Subagent | The output is large and you will not read it again; only the summary returns [1] |
| Fetch documentation or process a log file | Subagent | The same: the text stays in the subagent's context [1] |
| Research three modules that do not depend on each other | Several subagents at once | Each explores its area independently and Claude combines the findings [2] |
| A task that needs back-and-forth with you | Main conversation | Anthropic lists frequent back-and-forth and step-by-step refinement as main-conversation work [2] |
| Planning, implementing and testing one feature | Main conversation | The phases share a large amount of context [2] |
| A quick, targeted change | Main conversation | A subagent starts with a new context and may need time to gather context [2] |
| A side task that needs everything discussed so far | Fork | It inherits the history and reads the parent's cache [2] |
| A question about something already in the conversation | /btw |
It sees the full context, has no tools, and adds nothing to the history [2] |
| A task where tool access must be restricted | Subagent | The definition limits which tools it can use [2] |
The page on writing requests that read fewer files covers the main-conversation side of this choice, and the page on hooks that shorten tool output covers the other way to keep test output small: a hook is a script Claude Code runs by itself at a fixed point, such as before a command.
Our position
Delegate by the size of the output, and judge the result by the main conversation. A subagent is worth its start-up cost when the text it would otherwise put in your conversation is larger than its own start-up text and would be re-read on many later turns. A test run or a log search meets that condition. A one-line question or a change to one named function does not. When the task needs the history, fork; when it needs none of it, give it a prompt that stands alone and name the shape of the answer.
Once you delegate, the next question is which model the subagent runs on, because its requests are billed like any other. The next page, giving a subagent a smaller model, covers that setting, and what agent teams cost covers the more expensive option, an agent team, which is several Claude Code sessions working on one task at once.
Reveneau is an AI software development consultancy. All of its code is written by AI, and every change must pass an eval suite, a set of automated tests written from the specification, before release, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic, and every figure on this page is Anthropic's own statement about its own product, read on 4 October 2026.
Common questions
What is a subagent in Claude Code?
A subagent in Claude Code is a second copy of Claude that handles one task in its own context window, with its own system prompt, its own tool access and its own permissions. Anthropic's documentation, read on 4 October 2026, says the subagent works independently and returns results, and that it sends its own requests, which count toward the same usage limits as your main conversation. Claude delegates to a subagent when a task matches the subagent's description, or when you ask for one.
What comes back to the main conversation when a subagent finishes?
Only the subagent's final text response comes back to the main conversation, plus what Anthropic's context window page calls a small metadata trailer with token counts and duration. In that simulation, read on 4 October 2026, the subagent read 6,100 tokens of files and the main conversation received a 420-token result. The file contents, the tool calls and the subagent's reasoning stay in the subagent's own transcript, which is stored in a separate file.
Does a subagent see the files Claude already read in the main conversation?
No. A subagent starts with a fresh, isolated context window, and Anthropic's documentation, read on 4 October 2026, says it does not see your conversation history, the skills you have already invoked, or the files Claude has already read. Claude writes a delegation message that summarises the task, and the subagent works from that message. The one exception is a fork, which inherits the whole conversation at the moment it starts.
Does a subagent load my CLAUDE.md files?
Yes, most subagents load every level of the CLAUDE.md hierarchy that the main conversation loads, by Anthropic's documentation read on 4 October 2026. The built-in Explore and Plan subagents skip CLAUDE.md and the git status snapshot to keep research fast and inexpensive. A custom subagent can skip the user, project and local CLAUDE.md files with the field `omitClaudeMd: true` in its frontmatter, the settings block at the top of its file, which requires Claude Code v2.1.271 or later.
Where do I see how much of my usage went to subagents?
On a Pro, Max, Team or Enterprise plan, run `/usage` and read the attribution section, which shows recent usage attributed to skills, subagents, plugins and individual MCP servers, each as a percentage of the total. Anthropic's cost page, read on 4 October 2026, says the figures are computed from local session history on this machine and cover the last 24 hours or 7 days, switched with `d` and `w`. Usage from other devices is left out.
Why does a subagent's first request cost full price even when my main conversation's cache is still active?
A subagent's first request reads nothing from the cache because it starts its own conversation with its own system prompt and tool set, so its request prefix differs from the parent's and matches nothing the cache holds. Anthropic's prompt caching page, read on 4 October 2026, says the subagent then builds a cache of its own across its turns, and that the parent's cache is unaffected, because the subagent's call and result append to the end of the parent conversation.
What does a fork inherit from the main conversation?
A fork inherits the entire conversation at the moment it starts: the same system prompt, tools, model and message history as the main session, by Anthropic's documentation read on 4 October 2026. Its own tool calls still stay out of your conversation and only its final result returns. You start one yourself with `/subtask` followed by a task, which requires Claude Code v2.1.212 or later. A fork cannot start further forks.
What is a skill with context: fork?
A skill with `context: fork` in its frontmatter, the settings block at the top of its file, runs in a new subagent of the type named in its `agent` field, with the skill content as the subagent's prompt. Anthropic's documentation, read on 4 October 2026, says this subagent does not see your conversation history, so the skill's instructions have to stand on their own. Despite the name, it starts a new subagent; to hand over the history, fork the conversation instead.
Can a subagent be resumed instead of started again?
Yes. Anthropic's documentation, read on 4 October 2026, says each subagent invocation normally creates a new instance, and that Claude can resume a finished subagent by sending it a message with the `SendMessage` tool, using the agent ID or name. The resumed subagent keeps its full history, including earlier tool calls and results, and its first request can read the prompt cache that the original run built. The built-in Explore and Plan agents return no agent ID and cannot be resumed.
How many subagents can run at the same time in Claude Code?
By default, 20 subagents can run at the same time in a session. Anthropic's documentation, read on 4 October 2026, says that when 20 are running, starting another with the Agent tool fails with `Concurrent subagent limit reached`, and the error tells Claude not to retry. The limit is changed with `CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS`, requires Claude Code v2.1.217 or later, and sessions with the ultracode setting on are exempt. There is no limit on the total over a session.
Can a subagent start its own subagents?
Yes. By default a subagent can start subagents of its own, up to three layers below the main conversation, by Anthropic's documentation read on 4 October 2026. At the depth limit Claude Code withholds the Agent tool, so a subagent at that depth does the work itself and returns one summary. The limit is set with `CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH`, and a value of `1` turns nesting off. Each nested subagent sends its own requests, which count in your usage.
Do subagent descriptions use tokens in the main conversation?
Yes. Claude reads every subagent's `description` field to decide when to delegate, so those descriptions sit in the main conversation's context on every request. Anthropic's documentation, read on 4 October 2026, says that when the combined descriptions of your own subagents, excluding the built-in ones, exceed 15,000 tokens, Claude Code shows a warning at startup with the total token count and still loads every subagent. Its advice is to keep descriptions short and move detail into the subagent's system prompt, which loads only when that subagent runs.
References
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Create custom subagents (code.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Extend Claude with skills (code.claude.com), read 4 October 2026
- Anthropic, Explore the context window (code.claude.com), read 4 October 2026