The cost of agents and unattended runs in Claude Code
Every extra agent in Claude Code is an extra context window, and every context window sends its own billed requests. A subagent, an agent teammate, an agent in a dynamic workflow, a scheduled task firing while you are away, and a `claude -p` run started by a script all add requests that the main conversation never shows you. Anthropic's documentation, read on 4 October 2026, states the cost of each and names a control for each: the model an agent runs on, the spawn prompt it starts with, shutting it down, a spend cap, a time cap, and an output format that reports what a run cost. This guide collects those facts from Anthropic's own pages and ends with one real unattended job, Reveneau's own.
Published October 4, 2026. Editorial.
Key takeaways
- A subagent runs in its own context window and sends its own requests, which count toward the same usage limits as the main conversation, by Anthropic's documentation read on 4 October 2026.
- Anthropic's cost page puts agent team usage at 7 times a standard session when teammates run in plan mode, and gives four rules: Sonnet for teammates, small teams, focused spawn prompts, and shutting teammates down.
- A dynamic workflow run allows 1,000 agents in total and 16 at once by default, and shows a Large workflow warning past 25 agents or a projected 1.5 million tokens.
- A scheduled task fires on its interval even while the session is idle and sends the full context each time, and a recurring task expires 7 days after creation.
- --max-budget-usd caps the dollars a claude -p run spends, subagents included, in print mode only, and --output-format json makes the run report total_cost_usd.
Claude Code is Anthropic's coding tool: you type a request in a terminal, and an AI model reads files, runs commands and edits code for you. Everything the model reads and writes is counted in tokens, which are pieces of text. All of that text goes into the context window, which is the full set of text the model reads on each request, and Claude Code sends the full conversation with every request [1]. One conversation is one context window. The moment Claude Code starts a second copy of Claude, there are two.
That is the whole subject of this guide. A subagent, an agent teammate, an agent inside a dynamic workflow, a scheduled task firing while you are away, and a claude -p run started by a script all add requests that your main conversation never shows you. Anthropic's documentation, read on 4 October 2026, states what each one costs and names a control for each. This guide collects those statements, page by page, and ends with one real unattended job. Reveneau is an AI software development consultancy; all of its code is written by AI, and every change must pass an eval suite, a set of automated tests written from the specification, before release, so token use is a running cost of every Reveneau build, and the cost of agents is where that running cost multiplies.
Why every extra agent is an extra context window
Anthropic's cost page lists the reasons usage climbs in a long session, and three of them are agents. Every subagent, and every agent a dynamic workflow starts, sends its own requests in addition to the main conversation's [1]. "Each active teammate keeps consuming tokens until it exits." And a scheduled task "fires on its interval even while the session is idle, sending your full context each time" [1]. The subagent page says each subagent "runs in its own context window with a custom system prompt, specific tool access, and independent permissions", and that it "sends its own requests, which count toward the same usage limits as your main conversation" [2]. The agent teams page says teammates work "each in its own context window" [3].
Those sentences share one fact. Each agent starts with its own set of text to read, which Anthropic says includes CLAUDE.md, the instruction file you write for Claude, plus MCP servers, which are programs that connect Claude Code to outside tools, plus skills, which are packaged instruction sets, and then it works through its own turns, each of which sends everything it has read so far [1][2]. The main conversation receives a summary. The cost of producing that summary is paid in full somewhere else, and /usage on a Pro, Max, Team or Enterprise plan shows the share attributed to subagents, skills, plugins and MCP servers [1].
The controls follow from the same fact. If an agent is a context window, you can give it a cheaper model, give it less to read at the start, stop it when it is done, cap what the whole run may spend, cap how long the run may take, and make the run report what it spent. Each sub-page below takes one of those.
Subagents
A subagent is a second copy of Claude that takes one task into a separate context window and returns only its final text to the conversation that started it. Anthropic's cost page says to delegate test runs, documentation fetches and log processing to subagents so that "the verbose output stays in the subagent's context while only a summary returns to your main conversation", and in the same paragraph that "the subagent's own requests still draw on your usage" [1]. The saving is in the main conversation; the subagent's start-up and work are paid in full. By default 20 subagents can run at once in a session, and starting another fails with Concurrent subagent limit reached until the count drops; that limit requires Claude Code v2.1.217 or later [2].
The page on when a subagent saves tokens states the trade in Anthropic's own simulated figures, lists what a subagent loads before it reads a file, explains why its first request reads none of the parent's prompt cache, and gives the tasks Anthropic says belong in a subagent and the tasks that belong in the main conversation.
The page on giving a subagent a smaller model covers the model field in a subagent's definition. Anthropic's cost page says to specify model: haiku for simple subagent tasks, and warns that a switch to Opus in the session also applies to the subagents that inherit the session's model [1]. The page gives the order Claude Code uses to pick a subagent's model, the setting that puts every subagent on one model, and Anthropic's list prices per model with the date read.
Agent teams and workflows
An agent team is several Claude Code sessions working on one task: one lead and its teammates, each a separate instance with its own context window, sharing tasks and messaging each other [3]. Anthropic's cost page puts agent team usage at 7 times that of a standard session when teammates run in plan mode, "because each teammate maintains its own context window and runs as a separate Claude instance" [1]. Teams are experimental and off by default; CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 in settings.json or the environment turns them on [1][3]. Anthropic's four cost rules are to use Sonnet for teammates, keep teams small, keep spawn prompts focused (the spawn prompt is the message that starts an agent), and shut teammates down when their work is done, because each active teammate keeps consuming tokens until it exits or the session ends [1]. The page on what agent teams cost takes each rule in turn, lists what every teammate loads, and names the tasks Anthropic says are worth a team.
A dynamic workflow is a script Claude writes for a task you describe, which a program inside Claude Code (the runtime) executes in the background, and each agent the script starts is a subagent sending its own requests [4]. The runtime allows up to 16 concurrent agents by default, fewer when fewer CPUs are available, and 1,000 agents in total per run, a limit Anthropic says exists to prevent loops that never end [4]. When a run schedules more than 25 agents, or its projected token total passes 1.5 million, its progress line shows a Large workflow warning that names /workflows, where the run can be stopped [4]. The page on what a workflow costs and how to cap it covers Anthropic's cost section for workflows, the size guideline, how agents started together in one step (a fan-out) share a prompt cache, what a stopped run keeps, and where a cap can be set.
Scheduled tasks and sessions left open
A session you leave open can start turns on its own. Anthropic's scheduled tasks page describes /loop, which runs a prompt on repeat on a fixed interval or at an interval Claude chooses, and the cron (timer) tools behind it; a session can hold up to 50 scheduled tasks, a recurring task expires 7 days after creation, and CLAUDE_CODE_DISABLE_CRON=1 turns the scheduler off [7]. The cost page says each firing sends your full context, and lists beside it the goal check-in that Claude Code starts while background work keeps a goal waiting, capped at three idle check-ins per goal since v2.1.246, and the message from another of your sessions that arrives as a new turn unless crossSessionInbound is set to hold [1]. The page on loops and scheduled tasks gives, for each of these, how often it fires, what it sends and which setting stops it, with the Loops rows in /usage that show the count after the fact.
Unattended runs
An unattended run is claude -p started by a script, a scheduler or a CI job (an automated build job), with no person at the keyboard. Anthropic's non-interactive page says -p runs Claude Code non-interactively and exits with code 0 on success and a non-zero code on failure, that without --bare it loads the same hooks, MCP servers and CLAUDE.md an interactive session would, and that with --output-format json the response includes total_cost_usd and a per-model cost breakdown [5]. The CLI reference describes --max-budget-usd as the maximum dollar amount to spend on API calls before stopping, in print mode only (print mode is the -p form), with subagent spend counted; once spend reaches the cap, starting another subagent fails with Budget limit reached and running background subagents are stopped, with enforcement from Claude Code v2.1.217 [6]. --max-turns limits the number of agentic turns (rounds in which Claude responds and uses tools) and has no limit by default [6].
Three controls belong on every such command line, and the page on running Claude Code from a script with a spend cap covers them: the spend cap, the output format that reports cost, and --model with a full model name, because an alias resolves to the recommended version and changes over time. It also covers --allowedTools and the permission modes as cost controls, since in a -p run a tool call that would prompt a person is denied, and each denied attempt is a turn, and it reports what Anthropic's GitHub Actions and GitLab pages say about cost in CI.
A spend cap counts dollars, and a run that is waiting spends none. A stalled response stream, a development server that never exits, or a subagent that never reports keeps a run open while the dollar counter does not move. The page on timeouts, retries and cleanup for unattended runs gives Anthropic's time controls at three levels: per request, per command and process, and per run. It covers the 10 retries Claude Code makes on its own, CLAUDE_CODE_RETRY_WATCHDOG for a run that must wait out an outage, what happens to a background process when claude -p ends, and the three numbers a retry loop of your own needs: an attempt count, a backoff and a budget.
The worked example
Reveneau runs its own hourly unattended job: a scheduler starts claude -p once an hour to write a short news item for this site, with nobody reviewing the item before it is published. The page on an hourly unattended job and its measured cost describes that job from its scripts and logs, read on 4 October 2026: the chain of steps, the time caps and the three incidents that produced them, the durations it measured, the dollar figures one of its steps records through total_cost_usd, and the cost its main step never measured because it runs with the default text output. That page says plainly that one job is one job and that none of its figures is a statistic. It is there so a reader can see what measuring an unattended run looks like, including the gap.
Each agent, what it adds, and its control
| What runs | What it adds, by Anthropic's documentation read on 4 October 2026 | The control |
|---|---|---|
| Subagent | Its own context window and its own requests, which count toward the same usage limits [2] | model: haiku for simple tasks; a focused task message; CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS [1][2] |
| Agent teammate | A separate Claude Code instance with its own context window; 7 times a standard session in plan mode [1][3] | Sonnet for teammates, small teams, focused spawn prompts, shut down when done [1] |
| Workflow agent | One subagent per agent() call; up to 16 at once and 1,000 per run [4] |
The size guideline; the Large workflow warning; stopping the run from /workflows [4] |
| Scheduled task | A turn on every firing that sends the full context [1] | CronDelete, Esc for a self-paced loop, the 7-day expiry, CLAUDE_CODE_DISABLE_CRON=1 [7] |
| Goal check-in | A turn while background work keeps a goal waiting, at most three idle check-ins per goal [1] | CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 [1] |
| Cross-session message | A new turn when the session is idle [1] | crossSessionInbound: hold [1] |
claude -p run |
Everything an interactive session loads, with no one watching [5] | --max-budget-usd, --max-turns, --output-format json, --model, --allowedTools, --bare [5][6] |
Where this guide meets the other three
This guide is one of four on Claude Code cost, and it assumes the other three. The guide to reducing Claude Code token usage covers the single conversation: where its tokens go, /clear against /compact, the length of CLAUDE.md, and which model and effort level to pick. Everything there applies inside every agent this guide adds, because each agent is a conversation of its own. The guide to the Claude Code prompt cache covers the store of request text the service reuses at a lower price, which decides what an agent's second request costs after its first; a subagent's first request reads none of the parent's cache, and agents a workflow starts together share one. The guide to Claude Code costs for teams covers the account side: what Anthropic says a developer costs, where spend appears, and the spend limits an administrator sets, which meet every request whichever agent sent it. The --max-budget-usd flag appears there in one sentence and here in full.
Our position
Treat every agent as a context window you are paying to fill, and ask of each one what it will read before it does any work. Start with a subagent on a smaller model, and give it a task message that says what to return. Use an agent team or a workflow when the task needs parallel full sessions or dozens of agents, and cap them in the ways Anthropic's pages name before the first run. Cancel loops before you leave a session, and set crossSessionInbound to hold on a session you leave idle. Give every unattended run --output-format json first, so that it reports its cost, then a spend cap set from that record, then a time cap, because a cap chosen without a record is a guess. Then read the worked example, which is what one of our own runs looks like with those rules applied and one of them still missing.
Reveneau is independent of Anthropic. Every figure in this guide about Claude Code is Anthropic's own statement about its own product, read on 4 October 2026, with the source named beside it, and the one real job described here is described from its own logs.
Explore the guide
Subagents
When a subagent saves tokens in Claude Code, and when it adds them
A subagent saves tokens in your main conversation when the task produces output you will never read again: a test run, a log file, a documentation fetch. A subagent is a second copy of Claude with its own conversation, its own system prompt and its own tools, and only its summary comes back. Anthropic's documentation, read on 4 October 2026, says the subagent's own requests still count toward the same usage limits as the main conversation. So the saving is in the main conversation, and the subagent's start-up and work are paid in full. This page explains when that trade is worth it, what a subagent loads before it works, why its first request reads none of the parent's prompt cache, and how to prompt it.
How to give a Claude Code subagent a smaller model
To give a Claude Code subagent a smaller model, set the `model` field in its definition file, for example `model: haiku`, which Anthropic's cost page recommends for simple subagent tasks. A subagent is a second copy of Claude that works in its own context window and sends its own billed requests. Without a `model` field, a subagent inherits the main conversation's model, so switching the session to Opus moves those subagents to Opus too. To put every subagent on one model, set `CLAUDE_CODE_SUBAGENT_MODEL` together with `CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1`, which requires Claude Code v2.1.257 or later. On Anthropic's pricing page, read on 4 October 2026, Claude Haiku 4.5 is listed at $1 per million input tokens against $4 for Claude Opus 5.5.
Teams and workflows
What an agent team costs in Claude Code
An agent team in Claude Code costs one full session per teammate. Each teammate is a separate Claude Code instance with its own context window that loads your CLAUDE.md files, MCP servers and skills, and keeps using tokens until it exits. Anthropic's cost page, read on 4 October 2026, puts agent team usage at 7 times that of a standard session when teammates run in plan mode. Teams are experimental and off by default; the setting `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` turns them on. Anthropic gives four cost rules: use Sonnet for teammates, keep teams small, keep spawn prompts focused, and shut teammates down when their work is done. This page explains what multiplies with each teammate and names the tasks Anthropic says are worth it.
What a Claude Code workflow costs and how to cap it
A Claude Code workflow costs one subagent per agent the script starts, and a single run can start dozens to hundreds of them. A dynamic workflow is a JavaScript script that Claude writes for a task you describe and that Claude Code executes in the background; each `agent()` call in it starts a subagent that sends its own billed requests. Anthropic's documentation, read on 4 October 2026, caps a run at 1,000 agents, runs 16 at once by default, and shows a `Large workflow` warning when a run schedules more than 25 agents or its projected token total passes 1.5 million. The size guideline, `small`, `medium` or `large`, tells Claude how many agents to aim for, and the `/workflows` view shows each agent's tokens while it runs.
Scheduled and unattended
What a loop or a scheduled task costs in Claude Code while you are away
A loop or a scheduled task in Claude Code costs one full turn on every firing, and a turn sends the whole conversation to the model. A scheduled task is a prompt that Claude Code runs again by itself on a timer, with `/loop` or the cron tools. Anthropic's cost page, read on 4 October 2026, says a scheduled task fires on its interval even while the session is idle, sending your full context each time. The same page lists three more ways a session starts a turn with nobody typing: a goal check-in, a message from another of your sessions, and the short request behind a prompt suggestion. This page says how often each one fires, what Anthropic's documentation says it sends, and which setting stops it.
How to run Claude Code from a script with a spend cap
To run Claude Code from a script with a spend cap, call `claude -p` with your prompt and add `--max-budget-usd` with a dollar amount. `claude -p` is a run with no person at the keyboard: Claude Code reads the prompt, works, prints the result and exits. Anthropic's CLI reference, read on 4 October 2026, says `--max-budget-usd` is the maximum dollar amount to spend on API calls before stopping, works in print mode only, counts spend from subagents, and from Claude Code v2.1.217 fails any further subagent with `Budget limit reached` once the cap is reached. Add `--output-format json` so the run reports `total_cost_usd`, and `--model` with a full model name so the run stays on one model version. This page covers each flag, what it controls and its default.
Timeouts, retries and cleanup for an unattended Claude Code run
An unattended Claude Code run needs a time cap as well as a spend cap, because a run that is waiting spends no dollars and still holds its place in the schedule. `--max-budget-usd` counts money spent on API calls; a stalled response stream, a development server that never exits, or a subagent that never reports keeps the session open while that counter does not move. Anthropic's documentation, read on 4 October 2026, gives time controls at three levels: per request, through `API_TIMEOUT_MS` and the streaming watchdogs; per command, through the Bash timeouts and the 30-minute limit on background commands in unattended sessions; and per run, through `--max-turns`, the 10-minute wait for background work, and a job-level timeout in CI (automated build jobs). This page covers each, plus the retry rules.
Common questions
Why does every extra agent in Claude Code cost an extra context window?
Because each agent is a separate copy of Claude with its own conversation, and every request sends that conversation to the model. Anthropic's documentation, read on 4 October 2026, says a subagent runs in its own context window and sends its own requests, that each agent teammate is a separate Claude Code instance with its own context window, and that every agent a dynamic workflow starts sends its own requests in addition to the main conversation's. The main conversation shows you only the summary that returns.
What is the difference between a subagent, an agent team and a workflow in Claude Code?
A subagent is one extra copy of Claude that takes one task into its own context window and returns a summary to the conversation that started it. An agent team is several full Claude Code sessions, one lead and its teammates, each in its own context window, sharing tasks and messaging each other. A dynamic workflow is a script Claude writes that starts many subagents at once and returns one result. Anthropic's documentation, read on 4 October 2026, describes each on its own page.
Should I start with a subagent, an agent team or a workflow?
Start with a subagent, on a smaller model where the task is simple. It is the only one of the three that adds a single context window, and Anthropic's cost page, read on 4 October 2026, recommends `model: haiku` for simple subagent tasks. Anthropic puts agent team usage at 7 times a standard session when teammates run in plan mode, and a workflow run can start up to 1,000 agents. Use those when the task needs parallel full sessions or dozens of agents, and cap them when you do.
What is an unattended run in Claude Code?
An unattended run is a `claude -p` call started by a script, a scheduler or a CI job (an automated build job), with no person at the keyboard. Anthropic's documentation, read on 4 October 2026, says `-p` runs Claude Code non-interactively, exits with code 0 on success and a non-zero code on failure, and, without `--bare`, loads the same hooks, MCP servers and CLAUDE.md an interactive session would. Its caps have to be on the command line, because nobody is watching the token count.
Which two caps should every unattended Claude Code run have?
A spend cap and a time cap. The spend cap is `--max-budget-usd`, which Anthropic's CLI reference, read on 4 October 2026, describes as the maximum dollar amount to spend on API calls before stopping, in print mode only, with subagent spend counted. The time cap is yours to supply, because a run that is waiting on a stalled request or an open process spends no dollars: a job-level timeout in CI, or an attempt cap in your own script, beside `--max-turns`, which bounds turns and has no limit by default.
What is a spawn prompt in Claude Code?
The spawn prompt is the message that starts an agent, and everything in it enters that agent's context from its first request. Anthropic's cost page, read on 4 October 2026, says to keep spawn prompts focused, because teammates load CLAUDE.md, MCP servers and skills automatically and everything in the spawn prompt adds to their context from the start. A subagent's task message enters its context the same way, so a long spawn prompt is paid on every request the agent makes.
Where can a dollar cap be set when Claude Code runs agents?
In a `claude -p` run, with `--max-budget-usd`, which Anthropic's CLI reference, read on 4 October 2026, lists as print mode only and counts subagent spend; the Agent SDK has the same cap as `maxBudgetUsd`. In an interactive session the dollar controls are the ones an organisation sets: workspace spend limits on the Claude Console, spend limits in the admin console on a Teams or Enterprise plan, or a Claude apps gateway cap, which the guide to Claude Code costs for teams covers.
What is the first control to add to an existing unattended Claude Code job?
`--output-format json`, so the run reports what it cost. Anthropic's non-interactive page, read on 4 October 2026, says that with this format the response includes `total_cost_usd` and a per-model cost breakdown, so scripted callers can track spend without the usage dashboard. A job that runs with the default `text` format records no cost at all, and a dollar cap chosen without that record is a guess. Measure first, then set `--max-budget-usd` above the largest healthy run.
Does this guide report a cost figure for Reveneau's own hourly job?
The pillar page reports none. The worked example page describes Reveneau's own hourly unattended job from its scripts and logs, read on 4 October 2026, with the time caps it uses, the incidents that produced them, and the dollar figures its second step recorded through `total_cost_usd`. That page says plainly that one job is one job and that none of its figures is a statistic about Claude Code or about any other job. It also says which cost the job never measured, and why.
Which controls in this guide apply to an interactive session as well as a script?
The model an agent runs on, the spawn prompt, shutting a teammate down, the workflow size guideline, and the loop and goal check-in settings all apply in an interactive session, by Anthropic's documentation read on 4 October 2026. `--max-budget-usd` and `--max-turns` apply in print mode only, which means a `claude -p` run or the Agent SDK. The organisation-level spend limits apply to both, because they are set on the account side and meet every request.
References
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Create custom subagents (code.claude.com), read 4 October 2026
- Anthropic, Orchestrate teams of Claude Code sessions (code.claude.com), read 4 October 2026
- Anthropic, Orchestrate subagents at scale with dynamic workflows (code.claude.com), read 4 October 2026
- Anthropic, Run Claude Code programmatically (code.claude.com), read 4 October 2026
- Anthropic, CLI reference (code.claude.com), read 4 October 2026
- Anthropic, Run prompts on a schedule (code.claude.com), read 4 October 2026