The cost of agents and unattended runs in Claude Code / Teams and workflows
What a Claude Code workflow costs and how to cap it
A Claude Code workflow costs one subagent per agent the script starts, and a single run can start dozens to hundreds of them. A dynamic workflow is a JavaScript script that Claude writes for a task you describe and that Claude Code executes in the background; each `agent()` call in it starts a subagent that sends its own billed requests. Anthropic's documentation, read on 4 October 2026, caps a run at 1,000 agents, runs 16 at once by default, and shows a `Large workflow` warning when a run schedules more than 25 agents or its projected token total passes 1.5 million. The size guideline, `small`, `medium` or `large`, tells Claude how many agents to aim for, and the `/workflows` view shows each agent's tokens while it runs.
Published October 4, 2026. Editorial.
Key takeaways
- A dynamic workflow is a script Claude writes that starts many subagents at once; Anthropic's documentation, read on 4 October 2026, puts its scale at dozens to hundreds of agents per run.
- The runtime allows up to 16 concurrent agents by default, 4,096 items in one parallel() or pipeline() call, and 1,000 agents in total per run.
- A Large workflow warning appears when a run schedules more than 25 agents or its projected token total passes 1.5 million; it is advisory and does not pause the run.
- The size guideline is advice to Claude: small aims for fewer than 5 agents, medium for fewer than 10, large for fewer than 50, with medium the default and small the default on a Pro plan from v2.1.271.
- Agents in a fan-out (one step that starts many agents at once) that match on model, effort, agent type, tools, schema and directory share a prompt cache prefix, and Claude Code holds all but the first for up to 5 seconds so they read it.
Claude Code is Anthropic's coding tool: an AI model that reads files, runs commands and edits code from a request you type in a terminal. A subagent is a second copy of Claude that takes one task into its own context window, the full set of text a model reads on each request, and sends its own requests, billed in tokens. A dynamic workflow is the way Claude Code starts many subagents at once from a script, and this page is about what that costs and where the limits are.
It is part of the guide to the cost of agents and unattended runs in Claude Code. Every fact on it comes from Anthropic's documentation, read on 4 October 2026, and the cost section of Anthropic's workflows page is quoted in full where it matters.
What a dynamic workflow is
Anthropic's definition: "A dynamic workflow is a JavaScript script that orchestrates many subagents at once. Claude writes the script for the task you describe, and a runtime executes it in the background while your session stays responsive." [1] The runtime is the program inside Claude Code that runs the script. Dynamic workflows are available on all paid plans, with Anthropic API access, and on Amazon Bedrock, Google Cloud's Agent Platform and Microsoft Foundry; on a Pro plan you turn them on from the Dynamic workflows row in /config [1].
The difference from a subagent or an agent team is who holds the plan. With subagents and teams, Claude decides turn by turn what to start next, and every result goes into a context window. "A workflow script holds the loop, the branching, and the intermediate results itself, so Claude's context holds only the final answer." [1] Anthropic's comparison table puts the scale of subagents at "a few delegated tasks per turn", of agent teams at "a handful of long-running peers", and of workflows at "dozens to hundreds of agents per run" [1].
The script is plain JavaScript. "agent() spawns one subagent, pipeline() runs one per item in a list, and parallel() runs a set of agent tasks at the same time and waits for all of them." [1] Spawn is Anthropic's verb for starting an agent. Every run writes its script to a file under your session's directory in ~/.claude/projects/, so you can read it, compare it with an earlier run's script, or edit it and ask Claude to relaunch it [1]. A run you want to repeat can be saved as a command from /workflows with s [1].
You start one by including the keyword ultracode in a prompt, or by asking in your own words, for example "use a workflow" [1]. A second route turns it on for the whole session: /effort ultracode makes Claude plan a workflow for each substantive task, and a session started with claude --effort ultracode also sets the effort level, the setting for how much the model reasons, to xhigh; that flag requires Claude Code v2.1.203 or later [1].
The cost section, in Anthropic's words
The workflows page has a section headed "Cost". It opens by saying that a workflow starts many agents, so a single run can use more tokens than working through the same task in conversation, and that "Runs count toward your plan's usage and rate limits." [1]
It then gives the method: to gauge the spend before committing to a large task, run the workflow on a small part first, "one directory instead of the whole repo, or a narrow question instead of a broad one." It continues: "The /workflows view shows each agent's token usage as the run progresses, and you can stop the run there at any time, usually without losing completed work." [1] The runtime's agent caps limit how many agents a single run can start, which Anthropic says bounds the cost of a script that keeps starting agents, and "To keep runs to fewer agents, choose the small size guideline." [1]
Each agent a workflow starts is a subagent, so the previous pages apply to it. It sends its own requests, which the cost page lists among the reasons usage climbs: every subagent, and every agent a dynamic workflow starts, sends its own requests in addition to the main conversation's, and "The attribution breakdown shows the subagent share." [2] That breakdown is the attribution section of /usage on a Pro, Max, Team or Enterprise plan, which shows recent usage attributed to skills (packaged instructions for one task), subagents, plugins and MCP servers (programs that connect Claude Code to outside tools) as percentages [2]. Anthropic's cost page groups workflow agents with subagents in that sentence and names no separate row for workflows. The page on how to read /usage and /context covers the screen.
The cost of ultracode is stated separately. With it on, "A single request can turn into several workflows in a row: one to understand the code, one to make the change, and one to verify it. This applies to every task in the session, so each request uses more tokens and takes longer than the same request without a workflow. On a subscription plan those tokens draw on your usage limits, so a session with ultracode on reaches a session or weekly limit sooner than the same work with it off." [1]
The limits the runtime enforces
Anthropic's table of constraints gives the fixed limits, each with its reason [1]:
| Limit | Value | Anthropic's reason |
|---|---|---|
| Concurrent agents | Up to 16 by default, fewer when Claude Code has fewer CPUs available; CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS sets 1 to 256, which requires Claude Code v2.1.269 or later |
"Bounds local resource use" |
Items in one parallel() or pipeline() call |
Up to 4,096; a longer list is rejected with an error | "A silent cap would drop part of the workload without telling the script" |
| Agents per run | 1,000 in total | To stop a loop that keeps starting agents |
| Fan-out start delay (a fan-out is one step that starts many agents at once) | Agents that share the first agent's cache prefix start up to 5 seconds after it by default | So all but the first read the prefix the first agent cached |
Two other constraints shape the cost. A run pauses on its own only for agent permission prompts and a usage-limit wait, so there is no mid-run user input; for sign-off between stages, Anthropic says to run each stage as its own workflow [1]. And when an agent() call asks for structured output with a schema, which is a fixed shape the answer must match, a subagent whose output still fails the shape check after five attempts fails the call; MAX_STRUCTURED_OUTPUT_RETRIES changes the attempt count [1].
The concurrent subagent limit from the subagents page, 20 by default, applies to subagents Claude starts with the Agent tool; workflow agents and agent team teammates "follow their own limits instead" [4]. The page on when a subagent saves tokens covers that limit.
The warning and the size guideline
Two controls shape a run without being limits. Anthropic calls the first a warning: "When a workflow schedules more than 25 agents, or its projected token total passes 1.5 million, its progress line in the task panel below the input box shows a Large workflow warning. The warning points you to /workflows, where you can stop the run." [1] It is advisory: "it doesn't pause or limit the run" [1]. A size guideline you choose replaces the 25-agent threshold with its own agent count, and sessions with ultracode on show no warning, "because turning ultracode on already opts you in to large runs" [1].
The second is the size guideline, which "tells Claude how many agents to aim for when it writes a dynamic workflow. Claude Code sends the guideline to Claude as advice, not a cap, so a prompt that calls for a different scale still overrides it." [1] It requires Claude Code v2.1.202 or later. The values [1]:
| Value | Agent count Claude aims for |
|---|---|
unrestricted |
No guideline: Claude sizes the workflow to the task |
small |
Fewer than 5 agents |
medium |
Fewer than 10 agents |
large |
Fewer than 50 agents |
The default is medium, or small when you are signed in on a Pro plan with Claude Code v2.1.271 or later; v2.1.219 introduced the default, and earlier versions default to unrestricted [1]. Set it with the Dynamic workflow size setting in /config, with /config workflowSizeGuideline=small, or, on v2.1.219 and later, with the workflowSizeGuideline key in any settings file, which takes precedence over /config [1]. Changes take effect on the next prompt, and the runtime caps still apply regardless [1].
Ultracode also turns off two other checks. With it on, the session's concurrent subagent limit is not enforced for the subagents Claude starts with the Agent tool, and in auto permission mode you are not asked to approve the first workflow launch [1].
Caching in a fan-out: the 5-second hold
The prompt cache is a store, kept by the service that runs the model, of request text it has already processed, so unchanged text is billed at a lower rate. A subagent's first request normally reads none of the parent's cache, because its request starts differently [3]. In a workflow, sibling agents can read each other's. "Two agents that run with the same model, effort level, agent type, tools, output schema, and working directory build the same tools-and-system-prompt prefix, so an agent that starts after a matching sibling's response has begun reads that sibling's cache on its first request." [1]
Claude Code arranges that order on purpose. "When a fan-out starts several matching agents at once, Claude Code holds all but the first until the first agent's response begins, then releases the held agents together so their first requests read the shared prefix instead of each processing it uncached. Claude Code caps the hold at CLAUDE_CODE_WORKFLOW_PREFIX_STAGGER_MS milliseconds, 5000 by default. Set it to 0 to disable the hold." [1] The prompt caching page lists the same mechanism among the requests that can read an earlier request's prefix [3].
A workflow agent's requests fall outside the main conversation's cache lifetime bucket, so its cache holds for five minutes by default, including on a Claude subscription; subagentPromptCacheTtl set to 1h extends it, and the API bills one-hour cache writes at a higher rate [1]. The page on the five-minute and one-hour cache lifetimes owns that setting.
What a stopped run keeps
Stopping a run is the main cost control once it has started, and Anthropic documents what a relaunch reuses. "Claude Code replays the run in the order agents started, and each agent either returns its saved result or runs again" [1]:
- Completed agents return their saved result. The first agent whose prompt differs from the previous run, because the script was edited or an earlier agent returned something different, runs again, and so does every agent after it.
- Still running when you stopped: the agent starts over. Stopping the whole run counts no agent as failed.
- Failed agents run again, and so does every agent that started after them, even ones that completed. Stopping one agent alone, with
xin/workflows, counts as failing it.
Anthropic's example of the last rule: "If a script starts A, B, C, and D in that order and B fails, relaunching returns A from cache and runs B, C, and D again." [1] A relaunch works within the same session, carries over when you background the session or choose Move to background and exit, and works in a session resumed with claude --resume, because the saved results stay under the session's directory; a new session starts the workflow over as a new run [1].
An agent() call resolves to null if you stop it mid-run or it hits an unrecoverable API error, and pipeline() keeps each null in its results [1]. Since Claude Code v2.1.271, an agent that reaches your claude.ai usage limit pauses the run instead of failing, under four conditions: an interactive session signed in with a claude.ai subscription, autoContinueAtUsageLimit on, a reset within 24 hours, and a run that has not already waited twice [1]. In a claude -p run, which is a run started from a script with no person at the keyboard, in the Agent SDK, in a background session, or in a teammate session, the agent fails instead [1].
Where a cap can be set
Anthropic's documentation gives these places to bound a workflow's cost, from the widest to the narrowest:
| Control | What it does | Where it is set |
|---|---|---|
| Turn workflows off | Removes the bundled workflow commands, the ultracode keyword and the Ultracode toggle; a run in progress keeps going |
/config, "disableWorkflows": true in ~/.claude/settings.json, or CLAUDE_CODE_DISABLE_WORKFLOWS=1; for an organisation, managed settings or the Claude Code admin settings page [1] |
| Agents per run | 1,000 in total, fixed | The runtime [1] |
| Concurrent agents | 16 by default, 1 to 256 | CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS, v2.1.269 or later [1] |
| Size guideline | Advice on the agent count Claude aims for | /config workflowSizeGuideline=small or the settings key [1] |
| Model per stage | The model each agent runs on | The script, CLAUDE_CODE_SUBAGENT_MODEL, or your session model [1] |
| Stop the run | Ends the run, keeping completed results for a relaunch | /workflows, select the run, press x [1] |
The model row needs one more sentence. "Claude Code picks each workflow agent's model in the same order it uses for subagents. A model the script names for a stage counts as the per-invocation model in that order. When nothing else assigns one, the agent runs on your session's model." [1] Anthropic's advice is to check /model before a large run and to ask Claude, when you describe the task, to use a smaller model for stages that do not need the strongest one [1]. The page on giving a subagent a smaller model covers the order and the force variable, which the subagents page says applies to workflow agents too.
For a run with no person watching, claude -p accepts a dollar cap, --max-budget-usd, and the page on running Claude Code from a script with a spend cap covers what it counts and what happens when it is reached.
Workflow shapes and what each agent adds
Anthropic's example prompts show the common shapes. The second column restates how each example describes its agents, and the third column follows from the facts above: each agent is a subagent with its own context and its own requests, and matching siblings started within the hold share a cache prefix.
| Workflow shape, from Anthropic's examples | Agents the example describes | What each agent adds |
|---|---|---|
| Audit many files for one issue | One agent per file, then agents that verify the findings | One context and one set of requests per file; verifiers add a second pass [1] |
| Keep fixing until a check passes | A checker, a fixer, repeated until the check passes or two rounds make no progress | A new round of agents each time the check fails [1] |
| Migrate many files in parallel | A discovery agent, one agent per file in its own isolated copy, then verification | One agent per file plus one verifier per result [1] |
| Review every changed file and write one summary | One reviewer per file, then one agent that ranks and removes duplicates | One context per file plus a final agent that reads every finding [1] |
| Research a topic across sources | Readers in parallel, then a synthesis; /deep-research is the bundled form |
One agent per source, each reading its own cache prefix if it matches a sibling [1] |
| Find issues until the list stops growing | Rounds of searching until two rounds find nothing new | A full set of agents per round, with no fixed round count [1] |
The agent counts in that table are the shapes Anthropic describes, with no numbers attached, because the number of files or sources sets them. The Large workflow warning at 25 agents and the per-run limit of 1,000 are the two fixed points on that scale.
Our position
Run the small part first, as Anthropic says, and read the /workflows view before you start the full run. Reveneau recommends the small size guideline as the setting on every machine where workflows are new, and a named model for every stage that reads or verifies, since those stages are the ones that multiply by the file count. Keep /effort ultracode off except for the session that needs it, because it turns every substantive task into one or more workflows and removes the warning that would have told you so.
By Anthropic's comparison table, a workflow is the option that starts the most agents per run from a single interactive session. The next pages leave the interactive session: loops and scheduled tasks covers work that runs on a timer while nobody types, and what agent teams cost covers the other way to run many sessions at once.
Reveneau is an AI software development consultancy. All of its code is written by AI, and every change must pass an eval suite, a set of automated tests written from the specification, before release, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic, and every limit, setting and version number on this page is Anthropic's own statement about its own product, read on 4 October 2026.
Common questions
What is a dynamic workflow in Claude Code?
A dynamic workflow in Claude Code is a JavaScript script that starts and coordinates many subagents at once. Anthropic's documentation, read on 4 October 2026, says Claude writes the script for the task you describe, a runtime executes it in the background while your session stays responsive, and intermediate results stay in script variables so Claude's context holds only the final answer. Each `agent()` call starts one subagent, `pipeline()` runs one per item in a list, and `parallel()` runs a set at the same time.
How many agents can one workflow run?
Up to 1,000 agents in total per run, by Anthropic's documentation read on 4 October 2026. The runtime runs up to 16 agents at the same time by default, fewer when Claude Code has fewer CPUs available, and `CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS` changes that to a value from 1 to 256, which requires Claude Code v2.1.269 or later. A single `parallel()` or `pipeline()` call accepts up to 4,096 items, and the runtime rejects a longer list with an error.
What is the Large workflow warning?
The `Large workflow` warning is a note on a run's progress line in the task panel that appears when a workflow schedules more than 25 agents or its projected token total passes 1.5 million, by Anthropic's documentation read on 4 October 2026. It points you to `/workflows`, where you can stop the run. The warning is advisory and neither pauses nor limits the run. A size guideline you choose replaces the 25-agent threshold with its own count, and sessions with ultracode on show no warning.
What is the workflow size guideline?
The workflow size guideline tells Claude how many agents to aim for when it writes a dynamic workflow. Anthropic's documentation, read on 4 October 2026, says Claude Code sends it as advice and that a prompt calling for a different scale overrides it. The values are `unrestricted`, `small` for fewer than 5 agents, `medium` for fewer than 10, and `large` for fewer than 50. The default is `medium`, or `small` on a Pro plan with Claude Code v2.1.271 or later. Set it with `/config workflowSizeGuideline=small`.
How do I see how many tokens a workflow has used?
Run `/workflows`, select the run, and press Enter to open its progress view. Anthropic's documentation, read on 4 October 2026, says the view shows each phase with its agent count, token total and elapsed time, and that you can drill into a phase to see its agents and what each one found. The cost section says the view shows each agent's token usage as the run progresses, and that you can stop the run there at any time, usually without losing completed work.
Does stopping a workflow lose the work its agents finished?
Usually no. Anthropic's documentation, read on 4 October 2026, says that when Claude relaunches a stopped run with the same script, it replays the run in the order agents started, and each completed agent returns its saved result instead of running again. An agent that was still running when you stopped starts over. The relaunch works within the same session, in a backgrounded session, and in a session resumed with `claude --resume`; a new session has no earlier run and starts over.
Why does one failed workflow agent make agents that had finished run again?
Because a failed agent runs again on relaunch, and so does every agent that started after it, even ones that completed. Anthropic's documentation, read on 4 October 2026, gives the example: if a script starts A, B, C and D in that order and B fails, relaunching returns A from its saved result and runs B, C and D again. Stopping one agent alone, with `x` in `/workflows`, counts as failing it, while stopping the whole run counts no agent as failed.
What is the 5-second hold in a workflow fan-out?
When a fan-out starts several agents that match on model, effort level, agent type, tools, output schema and working directory, Claude Code holds all but the first until the first agent's response begins, then releases the held agents together so their first requests read the prompt cache prefix the first agent built. Anthropic's documentation, read on 4 October 2026, caps the hold at `CLAUDE_CODE_WORKFLOW_PREFIX_STAGGER_MS` milliseconds, 5000 by default, and a value of 0 turns the hold off.
Which model do workflow agents run on?
Claude Code picks each workflow agent's model in the same order it uses for subagents, by Anthropic's documentation read on 4 October 2026: a model the script names for a stage counts as the per-invocation model, and when nothing else assigns one, the agent runs on your session's model. Anthropic advises checking `/model` before a large run and asking Claude to use a smaller model for stages that do not need the strongest one. A model blocked by `availableModels` is substituted, with a warning in `/workflows`.
Does /effort ultracode use more tokens?
Yes. With ultracode on, Claude plans a workflow for each substantive task instead of waiting for you to ask, and Anthropic's documentation, read on 4 October 2026, says one request can become several workflows in a row, so each request uses more tokens and takes longer than the same request without a workflow. On a subscription plan a session with ultracode on reaches a session or weekly limit sooner. Ultracode also removes the Large workflow warning and the concurrent subagent limit.
How do I turn workflows off?
For yourself, toggle Dynamic workflows off in `/config`, set `"disableWorkflows": true` in `~/.claude/settings.json`, or set the environment variable `CLAUDE_CODE_DISABLE_WORKFLOWS=1`, which is read at startup. For a whole organisation, set `"disableWorkflows": true` in managed settings or use the toggle on the Claude Code admin settings page. Anthropic's documentation, read on 4 October 2026, says a run already in progress keeps going, and that turning workflows off also makes ultracode unavailable.
What happens when a workflow run reaches my usage limit?
On Claude Code v2.1.271 or later, the run pauses: the agents that reached the limit wait for the reset, no new agents start, and the run continues on its own shortly after the reset. Anthropic's documentation, read on 4 October 2026, lists four conditions: an interactive session signed in with a claude.ai subscription, `autoContinueAtUsageLimit` on, a reset within 24 hours, and a run that has not already waited twice. When one fails, the affected agent fails instead; on earlier versions the affected agents always fail.
References
- Anthropic, Orchestrate subagents at scale with dynamic workflows (code.claude.com), read 4 October 2026
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Create custom subagents (code.claude.com), read 4 October 2026