The cost of agents and unattended runs in Claude Code / Scheduled and unattended
How to run Claude Code from a script with a spend cap
To run Claude Code from a script with a spend cap, call `claude -p` with your prompt and add `--max-budget-usd` with a dollar amount. `claude -p` is a run with no person at the keyboard: Claude Code reads the prompt, works, prints the result and exits. Anthropic's CLI reference, read on 4 October 2026, says `--max-budget-usd` is the maximum dollar amount to spend on API calls before stopping, works in print mode only, counts spend from subagents, and from Claude Code v2.1.217 fails any further subagent with `Budget limit reached` once the cap is reached. Add `--output-format json` so the run reports `total_cost_usd`, and `--model` with a full model name so the run stays on one model version. This page covers each flag, what it controls and its default.
Published October 4, 2026. Editorial.
Key takeaways
- --max-budget-usd caps the dollars a claude -p run spends on API calls, works in print mode only, and counts subagent spend, by Anthropic's CLI reference read on 4 October 2026.
- Once spend reaches the cap, starting another subagent fails with Budget limit reached and Claude Code stops background subagents that are still running; this enforcement requires Claude Code v2.1.217 or later.
- With --output-format json the response includes total_cost_usd and a per-model cost breakdown, which Anthropic describes as client-side estimates that can differ from the bill.
- --model with a full model name such as claude-opus-5-5 pins the run to one version; an alias such as opus resolves to the recommended version and changes over time.
- Anthropic's GitHub Actions and GitLab pages both name --max-turns and a timeout on the workflow or job as the cost controls for a run in CI (automated build and test jobs).
Claude Code is Anthropic's coding tool: you type a request in a terminal, and an AI model reads files, runs commands and edits code for you. Everything the model reads and writes is counted in tokens, which are pieces of text, and each request sends the context window, the full set of text the model reads on that request. When a script starts Claude Code instead of a person, nobody watches the token count rise, nobody answers a permission prompt, and nobody presses Escape. The caps have to be in the command line.
This page is part of the guide to the cost of agents and unattended runs in Claude Code. It covers claude -p, the flag that caps its spend in dollars, the output format that reports what it spent, the flags that pin its model and limit its tools, and what Anthropic's pages on GitHub Actions and GitLab say about cost. Every fact is from Anthropic's documentation, read on 4 October 2026.
What claude -p is
claude -p, also written --print, runs Claude Code non-interactively: it takes the prompt from the command line or from standard input, works, prints the response and exits [1]. Anthropic calls this the Agent SDK's command-line form; the same tools, agent loop and context management are also available as Python and TypeScript packages [1]. It is the form a scheduler, a CI job (an automated build and test job) or a shell script calls, and it is the run this page means by unattended: a run with no person at the keyboard.
Two behaviours matter for a script. The process exits with code 0 on success and a non-zero code when the run fails, so a script can branch on the exit status; an invalid flag is reported to standard error before the run starts, and a failure inside the run, such as missing authentication, is printed as the result on standard output [1]. And without --bare, a -p run loads the same context an interactive session would, including the hooks in a project's .claude/settings.json and the servers in its .mcp.json, even in a folder you have never trusted, with no trust dialog and no per-server approval prompt [1]. Anthropic's own examples add --bare for CI and scripts [1].
--max-budget-usd: the spend cap
Anthropic's CLI reference describes the flag in one entry, read on 4 October 2026 [2]:
- It sets the maximum dollar amount to spend on API calls before stopping, and it applies in print mode only (print mode is the
-pform). The example isclaude -p --max-budget-usd 5.00 "query". - Spend from subagents counts toward the cap. A subagent is a second copy of Claude that takes one task into its own context window and sends its own billed requests; the page on when a subagent saves tokens covers them.
- When you return to a conversation with
--continueor--resume, totals restored from earlier runs do not count toward the cap. - Once spend reaches the cap, starting another subagent fails with
Budget limit reached, and Claude Code stops background subagents that are still running. These enforcement behaviours require Claude Code v2.1.217 or later.
Two more facts come from other pages. The cost page says that for a response billed at the 1.1 times data residency rate, Claude Code multiplies the list price of that response's tokens by 1.1 in the session cost figure, and that the multiplied figure also counts toward --max-budget-usd; before v2.1.239 Claude Code did not apply the multiplier, so the figure was lower than the bill [3]. Anthropic's pricing page, read on 4 October 2026, defines that rate as the 1.1 times multiplier charged on all token categories when a request asks for US-only inference on Claude 4.6 and later models [10]. And the Agent SDK's cost tracking page says that when the cap ends a run, the result has the subtype error_max_budget_usd, its usage field leaves out the response that crossed the budget, and total_cost_usd includes it [4]. The SDK option is maxBudgetUsd in TypeScript and max_budget_usd in Python, and it counts only the call's own spend [4].
The cap measures dollars at Claude Code's own estimate, covered next. The page on spend limits for an organisation, a group and a person covers the limits an administrator sets on the account side, which apply to interactive sessions too.
--output-format json: reading what the run cost
--output-format takes three values in print mode: text, the default, prints plain text; json prints structured JSON with the result, the session ID and metadata; stream-json prints newline-delimited JSON as the run goes [1]. Anthropic's non-interactive page says that with --output-format json the response payload includes total_cost_usd and a per-model cost breakdown, so scripted callers can track spend without consulting the usage dashboard. When the run continues an earlier conversation with --continue or --resume, it reports the conversation's whole total, earlier runs' spend included [1]. With stream-json, the last line is a result message with the final response text, the cost and the session metadata [1].
What the figure is, by Anthropic's cost tracking page [4]:
total_cost_usdand the SDK'scostUSDare client-side estimates. The client computes them locally from a price table bundled at build time, unless amodelPricingtable from managed settings is in effect. They can drift from the bill when pricing changes, when the installed version does not recognise a model, or when a billing rule applies that the client cannot model. Anthropic says to use them for development insight and budgeting, and to take authoritative billing from the Usage and Cost API or the Console's usage page.total_cost_usdincludes subagent requests. Theusagefield in the same result counts only the top-level agent loop, so it undercounts as soon as the run delegates.modelUsage(model_usagein Python) includes subagents and splits the total by model, and each entry'scostBasissays whether list price, a managed table or neither priced it, which requires Claude Code v2.1.246 or later.- Success and error results both include the fields. A session crash can leave them zeroed in the final result, so recover totals from the result before the crash.
- Independent calls each report their own cost, so a script adds them. Calls that resume one session already include the session's earlier spend, so read the latest result and do not sum. Before v2.1.277, a resumed
claude -pcall started its totals at zero.
The practical consequence: a job that calls claude -p with the default text format records no cost at all. Reveneau's own hourly job did this for its main step, and the page on an hourly unattended job and its measured cost says what that gap looks like and how the second step avoided it.
Pin the model with --model
--model sets the model for the session, with an alias such as sonnet, opus, haiku or fable, or a model's full name, and it overrides the model setting and the ANTHROPIC_MODEL variable [2]. Without it, Claude Code resolves the model in Anthropic's priority order: --model, then ANTHROPIC_MODEL, then the model field in a settings file, which is also where /model saves your choice, then ANTHROPIC_DEFAULT_MODEL [8]. A script that passes no --model, on a machine where no environment variable sets one, therefore runs on whatever a person last saved in their settings.
An alias and a full name behave differently over time. Anthropic says aliases point to the recommended version for your provider and update over time, and that to pin a specific version you use the full model name, for example claude-opus-5-5, or set the matching ANTHROPIC_DEFAULT_OPUS_MODEL variable [8]. On the Anthropic API, opus resolves to Opus 5.5 and sonnet to Sonnet 5.5; the resolution differs on Amazon Bedrock, Google Cloud's Agent Platform and Microsoft Foundry [8]. When the configured model is one the account cannot use, the run fails with There's an issue with the selected model, and for -p Anthropic's remedy is to pass --model with a valid alias or ID or set ANTHROPIC_MODEL [9]. The page on giving a subagent a smaller model covers the model a run's subagents use, which is separate from this flag.
Fewer prompts, fewer wasted turns
In a -p run, a tool call that would prompt a person has nobody to ask. Anthropic's permission modes page says that in auto mode a non-interactive run without a --permission-prompt-tool has no prompt to fall back to, and that when repeated blocks reach a threshold the action does not run and Claude keeps working; Claude Code does not stop the run [7]. On our reading, each refused attempt is model output and a further request, so the tools a run is allowed to use decide how many turns it spends being refused.
Three controls shape that, by Anthropic's non-interactive page [1]:
--allowedTools, also written--allowed-tools, names tools that run without prompting, using permission rule syntax.ReadandEditallow file reads and edits;Bash(git diff *)allows any command starting withgit diff, and the space before the*matters, becauseBash(git diff*)would also matchgit diff-index[1][2].--toolsis the different flag that restricts which tools exist at all [2].--permission-modesets the baseline.acceptEditswrites files without prompting and auto-approves common filesystem commands such asmkdir,touch,mvandcp; other shell commands and network requests still need an--allowedToolsentry or apermissions.allowrule.dontAskdenies every call that would otherwise prompt, which Anthropic suggests for locked-down CI runs.autohas a separate reviewing model, which Anthropic calls the classifier, check most actions. A run that sets no mode starts indefaultwhen it fetches feature flags, and inautoon v2.1.285 or later when it does not, so Anthropic's advice is to pass the mode you want [1][7].--permission-prompts none, which requires Claude Code v2.1.259 or later, is for a run where nobody can answer. Anything that would prompt is denied unless aPermissionRequesthook allows it, Claude is told that nobody can approve the request and not to retry it, and the run continues. Claude Code also removes the tools that need an answer from a person, such asAskUserQuestion[1].
The reason these are cost controls is the retry. A denied call with no instruction against retrying can be tried again in a new turn. Reveneau recommends naming every tool the task needs up front and passing --permission-prompts none, so that a refused call ends with one refusal instead of several.
--max-turns and --bare
--max-turns limits the number of agentic turns in a print-mode run and exits with an error when the limit is reached. There is no limit by default [2]. Anthropic's cost tracking page describes one round of that loop as: Claude responds, uses tools, gets results, and responds again [4]. The flag bounds how many of those rounds a run can take. Each round can send any amount of context, so the turn cap belongs beside the dollar cap.
--bare reduces start-up work by skipping auto-discovery of hooks, skills, custom commands, subagents, installed plugins, MCP servers, auto memory and CLAUDE.md, the instruction file you write for Claude [1]. An MCP server is a program that connects Claude Code to an outside tool or data source, and in bare mode only servers passed on the command line connect. Claude keeps the Bash, file read and file edit tools; system prompt additions, settings, MCP servers, custom agents and plugins are passed by flag. Bare mode never reads the sign-in credentials of a Claude subscription login (OAuth), so for the Anthropic API you set ANTHROPIC_API_KEY [1]. Anthropic calls --bare the recommended mode for scripted and SDK calls and says it will become the default for -p in a future release [1]. For cost, the point is that nothing a teammate left in ~/.claude enters the run's context.
The Agent SDK, GitHub Actions and GitLab
The Agent SDK's cost tracking page is Anthropic's reference for reading cost from a programmatic run. Beyond the fields above, it says per-step output_tokens on assistant messages is a placeholder and the real output count is on the result message, that parallel tool calls share one message ID and must be counted once, and that cache_creation_input_tokens and cache_read_input_tokens should be tracked apart from input_tokens to see what the prompt cache saved [4]. The prompt cache is the service's store of request text it has already processed, billed at a lower rate on reuse; the page on how to check your cache hit rate covers the same fields in an interactive session.
Anthropic's GitHub Actions page says each run consumes two resources, GitHub Actions minutes and API tokens, and that an OAuth token makes runs draw on a Claude subscription instead of API billing. Its cost list is: write specific @claude requests so Claude needs fewer turns, use issue templates for context, keep CLAUDE.md concise because Claude reads it on every run, set --max-turns in claude_args, set workflow-level timeouts, and use GitHub's concurrency controls to limit parallel runs [5]. claude_args accepts any CLI argument; Anthropic's example is --max-turns 5 --model claude-sonnet-5 --mcp-config /path/to/config.json [5]. The GitLab page lists runner time and API tokens as the two costs and names --max-turns, the job-level timeout keyword, for example timeout: 30m, and limited concurrency as the controls [6]. Neither page mentions --max-budget-usd; the flag is a CLI argument and claude_args passes any CLI argument, so Reveneau recommends adding it there as well.
Each flag, what it controls, and its default
| Flag or setting | What it controls, by Anthropic's documentation | Default |
|---|---|---|
-p, --print |
Runs non-interactively and exits | Interactive without it [1] |
--max-budget-usd |
Dollars spent on API calls before stopping, subagents included; print mode only; enforcement from v2.1.217 | No cap [2] |
--max-turns |
Number of agentic turns; exits with an error at the limit; print mode only | No limit [2] |
--output-format |
text, json with total_cost_usd and per-model cost, or stream-json |
text [1] |
--model |
The session's model, by alias or full name; overrides the model setting and ANTHROPIC_MODEL |
ANTHROPIC_MODEL, then the settings file's model, then ANTHROPIC_DEFAULT_MODEL [2][8] |
--allowedTools |
Tools that run without a prompt, in permission rule syntax | None listed; the permission mode decides what prompts or is denied [1] |
--permission-mode |
The run's baseline: default, acceptEdits, plan, auto, dontAsk, bypassPermissions |
default when feature flags are fetched; auto on v2.1.285 or later when they are not [7] |
--permission-prompts none |
Denies anything that would prompt and tells Claude not to retry; v2.1.259 or later | Prompts wait for a host, or are denied in a run with none [1] |
--bare |
Skips hooks, skills, commands, subagents, plugins, MCP servers, auto memory and CLAUDE.md | Everything an interactive session would load [1] |
maxBudgetUsd / max_budget_usd |
The SDK's form of the dollar cap; counts the call's own spend | No cap [4] |
Our position
Every unattended call to Claude Code should have four things on its command line: --max-budget-usd with a figure you chose, --output-format json so the run reports total_cost_usd, --model with a full model name, and an explicit permission mode with the tools the task needs in --allowedTools. Add --max-turns as the second cap and --bare when the run should be the same on every machine. Pick the dollar figure from the run's own history: read total_cost_usd for a week of runs, then set the cap above the largest healthy run, so that a normal run never meets it and an abnormal one does. Anthropic's example value is $5.00; the right value for your job is in your logs.
A dollar cap bounds spend only, and a run that waits on a stalled request or an open process holds its place in the schedule while spending nothing. The next page, timeouts, retries and cleanup for unattended runs, covers the time side.
Reveneau is an AI software development consultancy. All of its code is written by AI, and every change must pass an eval suite, a set of automated tests written from the specification, before release, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic, and every figure on this page is Anthropic's own statement about its own product, read on 4 October 2026.
Common questions
Does --max-budget-usd work in an interactive Claude Code session?
No. Anthropic's CLI reference, read on 4 October 2026, lists `--max-budget-usd` as print mode only, which means it applies to a `claude -p` run and to the Agent SDK's `maxBudgetUsd` or `max_budget_usd` option. In an interactive session the dollar controls are the ones an organisation sets: workspace spend limits on the Claude Console, spend limits in the claude.ai admin console on a Teams or Enterprise plan, or a Claude apps gateway cap. The flag takes a dollar amount, for example `--max-budget-usd 5.00`.
Does subagent spend count toward --max-budget-usd?
Yes. Anthropic's CLI reference, read on 4 October 2026, says spend from subagents counts toward the cap, and the Agent SDK's cost tracking page says the `total_cost_usd` field, which the cap is measured against, includes subagent requests alongside the top-level loop. The `usage` field in the same result excludes subagent activity, so a script that compares `usage` with the cap undercounts as soon as the run delegates work.
What happens when a claude -p run reaches its --max-budget-usd cap?
Claude Code stops spending on API calls. Anthropic's CLI reference, read on 4 October 2026, states two enforcement behaviours: starting another subagent fails with `Budget limit reached`, and Claude Code stops background subagents that are still running. Both require Claude Code v2.1.217 or later. The Agent SDK page adds that the result then has the subtype `error_max_budget_usd`, and that its `usage` leaves out the response that crossed the budget while `total_cost_usd` includes it.
Does spend from an earlier run count against the budget when I resume with --continue?
No. Anthropic's CLI reference, read on 4 October 2026, says that when you return to a conversation with `--continue` or `--resume`, totals restored from earlier runs do not count toward `--max-budget-usd`. The reported `total_cost_usd` does include them, because a resumed run reports the conversation's whole total, so the budget and the reported figure measure different things on a resumed run: the budget counts this call's own spend.
Where do I read the cost of a claude -p run?
Pass `--output-format json` and read `total_cost_usd` from the response, which Anthropic's non-interactive page, read on 4 October 2026, says is included together with a per-model cost breakdown so that scripted callers can track spend without the usage dashboard. With `--output-format stream-json`, the last line is a `result` message that holds the final text, the cost and the session metadata. With the default `text` format the run prints the answer alone and records no cost figure.
Is total_cost_usd from --output-format json an exact bill?
No. Anthropic's Agent SDK cost tracking page, read on 4 October 2026, calls `total_cost_usd` a client-side estimate computed locally from a price table bundled at build time, unless a `modelPricing` table is in effect. It can drift from the bill when pricing changes, when the installed version does not recognise a model, or when a billing rule applies that the client cannot model. Anthropic says to use it for development insight and budgeting and to take authoritative figures from the Usage and Cost API or the Console.
Should I pass an alias or a full model name to --model in a script?
Pass a full model name when the run must stay on one version. Anthropic's model configuration page, read on 4 October 2026, says aliases such as `opus` and `sonnet` resolve to the recommended version for your provider and update over time, and that to pin a specific version you use the full model name, for example `claude-opus-5-5`, or set the matching `ANTHROPIC_DEFAULT_OPUS_MODEL` variable. Either way, `--model` overrides the `model` setting and `ANTHROPIC_MODEL` for that run.
Which permission mode does claude -p start in when I set none?
In a session that fetches feature flags, a `claude -p` run starts in `default`, the mode the interface labels Manual, by Anthropic's permission modes page read on 4 October 2026. In a session that does not fetch them, such as one on a third-party provider or with telemetry off, it starts in `auto` on Claude Code v2.1.285 or later and `default` on earlier versions. Anthropic's advice is to pass `--permission-mode` so the run starts in the mode you intend.
What does --permission-prompts none do in an unattended run?
It tells Claude Code that nobody can answer a permission prompt. Anthropic's non-interactive page, read on 4 October 2026, says anything that would prompt is denied unless a `PermissionRequest` hook allows it, Claude is told that nobody can approve the request and not to retry it, and the run continues. Claude Code also removes tools that need a person, such as `AskUserQuestion`. In a `-p` run with no permission host the flag adds the instruction not to retry. It requires Claude Code v2.1.259 or later.
What does --bare change about a claude -p run?
`--bare` skips auto-discovery of hooks, skills, custom commands, subagents, installed plugins, MCP servers, auto memory and CLAUDE.md, so the run starts faster and loads the same context on every machine, by Anthropic's non-interactive page read on 4 October 2026. Claude keeps the Bash, file read and file edit tools; anything else is passed by flag. Bare mode never reads the credentials of a subscription login (OAuth), so set `ANTHROPIC_API_KEY` for the Anthropic API. Anthropic calls it the recommended mode for scripted calls.
How does --max-turns differ from --max-budget-usd?
`--max-turns` limits the number of agentic turns (rounds in which Claude responds and uses tools) in a print-mode run and exits with an error when the limit is reached, with no limit by default, by Anthropic's CLI reference read on 4 October 2026. `--max-budget-usd` limits dollars. A turn can be cheap or expensive depending on how much context it sends, so the two caps bound different things, and Anthropic's GitHub Actions and GitLab pages name `--max-turns` together with a job timeout as the controls for a CI run.
Which cost controls does Anthropic list for the Claude Code GitHub Action?
Anthropic's GitHub Actions page, read on 4 October 2026, lists six: write specific `@claude` requests so Claude needs fewer turns, use issue templates to provide context up front, keep `CLAUDE.md` concise because Claude reads it on every run, set `--max-turns` in `claude_args`, set workflow-level timeouts, and use GitHub's concurrency controls to limit parallel runs. It also says each run consumes GitHub Actions minutes as well as API tokens, and that an OAuth token makes runs draw on a Claude subscription.
References
- Anthropic, Run Claude Code programmatically (code.claude.com), read 4 October 2026
- Anthropic, CLI reference (code.claude.com), read 4 October 2026
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Track cost and usage, Agent SDK (code.claude.com), read 4 October 2026
- Anthropic, Claude Code GitHub Actions (code.claude.com), read 4 October 2026
- Anthropic, Claude Code GitLab CI/CD (code.claude.com), read 4 October 2026
- Anthropic, Choose a permission mode (code.claude.com), read 4 October 2026
- Anthropic, Model configuration (code.claude.com), read 4 October 2026
- Anthropic, Error reference (code.claude.com), read 4 October 2026
- Anthropic, Pricing (platform.claude.com), read 4 October 2026
More in Scheduled and unattended
What a loop or a scheduled task costs in Claude Code while you are away
A loop or a scheduled task in Claude Code costs one full turn on every firing, and a turn sends the whole conversation to the model. A scheduled task is a prompt that Claude Code runs again by itself on a timer, with `/loop` or the cron tools. Anthropic's cost page, read on 4 October 2026, says a scheduled task fires on its interval even while the session is idle, sending your full context each time. The same page lists three more ways a session starts a turn with nobody typing: a goal check-in, a message from another of your sessions, and the short request behind a prompt suggestion. This page says how often each one fires, what Anthropic's documentation says it sends, and which setting stops it.
Timeouts, retries and cleanup for an unattended Claude Code run
An unattended Claude Code run needs a time cap as well as a spend cap, because a run that is waiting spends no dollars and still holds its place in the schedule. `--max-budget-usd` counts money spent on API calls; a stalled response stream, a development server that never exits, or a subagent that never reports keeps the session open while that counter does not move. Anthropic's documentation, read on 4 October 2026, gives time controls at three levels: per request, through `API_TIMEOUT_MS` and the streaming watchdogs; per command, through the Bash timeouts and the 30-minute limit on background commands in unattended sessions; and per run, through `--max-turns`, the 10-minute wait for background work, and a job-level timeout in CI (automated build jobs). This page covers each, plus the retry rules.