The cost of agents and unattended runs in Claude Code / Scheduled and unattended
Timeouts, retries and cleanup for an unattended Claude Code run
An unattended Claude Code run needs a time cap as well as a spend cap, because a run that is waiting spends no dollars and still holds its place in the schedule. `--max-budget-usd` counts money spent on API calls; a stalled response stream, a development server that never exits, or a subagent that never reports keeps the session open while that counter does not move. Anthropic's documentation, read on 4 October 2026, gives time controls at three levels: per request, through `API_TIMEOUT_MS` and the streaming watchdogs; per command, through the Bash timeouts and the 30-minute limit on background commands in unattended sessions; and per run, through `--max-turns`, the 10-minute wait for background work, and a job-level timeout in CI (automated build jobs). This page covers each, plus the retry rules.
Published October 4, 2026. Editorial.
Key takeaways
- Claude Code retries transient API failures up to 10 times with exponential backoff; CLAUDE_CODE_MAX_RETRIES changes the count, capped at 15 as of v2.1.186, and Anthropic says to lower it to surface failures faster in scripts.
- In a claude -p run, a background shell task is terminated after the final result, which Anthropic puts at five seconds later, while a background subagent or workflow is waited for up to 10 idle minutes by default, set by CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS.
- In an unattended session, a background command gets 30 minutes, or the timeout Claude passes up to 2 hours, and is then stopped; this limit requires Claude Code v2.1.285 or later.
- CLAUDE_CODE_RETRY_WATCHDOG=1 retries 429 and 529 capacity errors without limit for unattended sessions, with backoff of up to 5 minutes, and fails at once on a 429 that reports a spend limit.
- A time budget alone does not bound retry attempts: Reveneau's own job ran 14 attempts in 71 seconds in a test on 12 September 2026 before an attempt count and a backoff were added.
Claude Code is Anthropic's coding tool: you type a request in a terminal, and an AI model reads files, runs commands and edits code for you. Everything the model reads and writes is counted in tokens, which are pieces of text, and each request sends the context window, the full set of text the model reads on that request. An unattended run is a claude -p call started by a script or a scheduler, with no person at the keyboard. The previous page gave it a spend cap. This page gives it a time cap.
This page is part of the guide to the cost of agents and unattended runs in Claude Code. It takes its facts about Claude Code from Anthropic's documentation, read on 4 October 2026, and its lessons about what goes wrong from Reveneau's own hourly unattended job, read from that job's scripts and logs on the same day. The full account of that job is on the page about an hourly unattended job and its measured cost; this page uses it only where a fact from it explains a control.
Why a spend cap needs a time cap beside it
--max-budget-usd is, in Anthropic's words, the maximum dollar amount to spend on API calls before stopping [2]. It counts money. A run that is waiting spends none: a response stream that has stopped delivering data, a background process that never exits, a subagent that never reports back. While the counter does not move, the session stays open, its context stays in memory, and if your scheduler takes a lock so that two runs never overlap, every later run waits too. Any turn the open session does take re-sends the full conversation, because Anthropic's cost page says Claude Code sends your full conversation with every request [4].
Reveneau's own hourly job gave an example of this, by its logs read on 4 October 2026. On 10 September 2026 the 17:07 run stayed open for 5 hours 47 minutes and blocked the next five hourly runs. The writer had finished its item by 17:21. The skill had asked it to confirm the new page rendered; no server was running, so it started a development server, which never exits, and the run waited on that server with the finished item unpublished. A dollar cap would have changed nothing about those hours, because the model was no longer being called.
Anthropic's documentation gives time controls at three levels, and the rest of this page takes them in turn: per request, per command and process, and per run. The whole-run cap for a job in CI, an automated build and test system, is the one the GitHub Actions and GitLab pages name: workflow-level timeouts, and GitLab's job-level timeout keyword, for example timeout: 30m [7][8]. In a script you run from your own scheduler, that outer cap is yours to write, and the last section says what it needs.
Processes a run starts, and who stops them
A claude -p run can start processes that outlive its last model request. Anthropic's non-interactive page says what happens to each kind [1]:
- A background Bash task, such as a development server or a watch build, is terminated after Claude has returned its final result and standard input has closed. Anthropic puts the grace period at five seconds, so a task that finishes just after the result can still deliver its output.
- A background subagent or workflow keeps the run open until the work completes, because its result is part of the final output. A subagent is a second copy of Claude that takes one task into its own context window; a dynamic workflow is a script Claude writes that starts many subagents. By default the wait ends after 10 minutes of continuous idle waiting, at which point Claude Code stops whatever is still running and drops its partial result.
CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MSchanges the limit,0waits without one, idle waiting starts over each time Claude takes a turn to handle a background result, and the variable requires Claude Code v2.1.182 or later [1][5]. - A Monitor watch, which is a background script that feeds each output line back to Claude, is waited for until it times out or the ten-minute cap ends the wait, whichever comes first. A watch times out five minutes after Claude starts it by default [1].
Anthropic's tools reference adds a limit on the commands themselves. In a session that runs unattended, such as a -p run, an Agent SDK application, a CI job or a cloud session, a background Bash or PowerShell command gets 30 minutes, or the timeout Claude passes with run_in_background, up to a maximum of 2 hours, counted from the moment it enters the background. Claude Code then stops it and tells Claude why, with the notice Background command "<description>" was stopped after reaching its background time limit. The limit requires Claude Code v2.1.285 or later, and before v2.1.288 it applied in every session. BASH_DEFAULT_TIMEOUT_MS above 1800000 replaces the 30-minute default, and BASH_MAX_TIMEOUT_MS above 7200000 raises the 2-hour maximum [6]. A foreground command has a default timeout of two minutes and an upper limit of ten minutes; at its timeout Claude Code moves it to the background instead of stopping it, unless the command starts with sleep, and in bare mode or with CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 it stops instead [6].
Stopping the run itself is covered too. A claude -p process stopped with SIGTERM (the standard stop signal) exits with code 143, leaves the turn in progress unfinished, terminates the process tree of any Bash command still running, runs SessionEnd hooks, and starts no new tool call or model request while exiting. SIGINT (the interrupt signal, which Ctrl+C sends) ends the turn instead [1]. For a subagent that stops making progress, CLAUDE_ASYNC_AGENT_STALL_TIMEOUT_MS, 600000 by default, aborts the subagent and reports the stall to the parent when no streaming progress arrives within the window [5].
Those rules cover what Claude Code itself starts and can see. Reveneau's hang shows the gap outside them, by the job's logs read on 4 October 2026: the development server had inherited the writer's output pipe (the channel that carried the writer's output to the runner), and the runner script was waiting for that pipe to close, which happens only when every process holding it has exited. The writer tried to stop the server with pkill, which was absent from its allowed tools, so the call was refused, and it wrote "That's OK; the process will get cleaned up" and stopped. Four changes followed: the writer's output goes to a file instead of a pipe; the runner ends the writer's whole process group (the writer and every process it started) after it exits, on every path including a clean one; a per-attempt time cap; and the runner starts and stops the server used to check the page itself and tells the writer never to start one. kill and pkill stayed out of the writer's allowed tools on purpose. The runner owns process lifetime, and Reveneau recommends the same split for any unattended job: the process that starts Claude Code also cleans up after it.
Retries Claude Code makes on its own
Anthropic's error reference says Claude Code retries transient failures up to 10 times with exponential backoff, which means the wait between attempts grows each time, before showing an error, and that by the time you see an error, the retries that apply have already happened [3]. The retried failures include server errors, overloaded responses and request timeouts that arrive before any of the response has streamed, dropped connections, and temporary 429 rate-limit responses. A failure that arrives after Claude has completed a block of text or a tool call is handled differently: Claude Code keeps what Claude completed, runs the finished tool calls, and continues the turn, because re-running the request could execute the same tool calls twice [3]. The Agent SDK's cost tracking page adds the cost rule: when a conversation fails midway, you still consumed tokens up to the point of failure, and both success and error results include total_cost_usd [9].
Three variables tune this, by the error reference and the environment variable reference [3][5]:
CLAUDE_CODE_MAX_RETRIES, default 10, sets the retry count. It is capped at 15 as of v2.1.186. Anthropic's advice for scripts is to lower it to surface failures faster.CLAUDE_CODE_RETRY_WATCHDOG=1is for unattended sessions; Anthropic's examples are eval harnesses (automated test runs), CI jobs (automated build jobs) and remote workers. It retries 429 and 529 capacity errors indefinitely instead of failing afterCLAUDE_CODE_MAX_RETRIESattempts, backing off up to 5 minutes between attempts, or until the limit resets when the response names a reset time. It fails at once when a standard-speed request gets a 429 that reports a spend limit or exhausted usage credits; before v2.1.239 it retried those indefinitely. On v2.1.199 or later it also raises the default retry count for other transient errors to 300, which Anthropic puts at three hours of backoff, and removes the cap of 15 onCLAUDE_CODE_MAX_RETRIES. It requires v2.1.186 or later.API_TIMEOUT_MS, default 600000, is the per-request timeout in milliseconds, 10 minutes. Anthropic says to raise it on slow networks or through a proxy, and warns that values above 2147483647 overflow the timer and make requests fail at once.
In a stream-json run you can watch the retries. Claude Code emits a system/api_retry event before each retry, with attempt, max_retries, retry_delay_ms, the HTTP status, and an error category such as rate_limit, overloaded, server_error or max_output_tokens [1].
When the API stalls
Anthropic's error reference separates three stalls, and each ends differently [3]:
- No response headers. When a streaming request gets no response headers within the first-byte deadline, Claude Code aborts it instead of waiting the full
API_TIMEOUT_MS, sends it again at most once if the retry budget allows, and ends the turn withNo response from APIif the retry goes unanswered too. The message shows how long each attempt waited. The retry waits one second less thanAPI_TIMEOUT_MS, which is 599 seconds with the default value (our arithmetic on Anthropic's 600000 milliseconds). Before v2.1.242 Claude Code waited the full timeout before failing. - Headers arrived, then nothing. When the headers have arrived but none of the response has, or Claude has finished thinking but started no text or tool call, Claude Code aborts the stalled connection and re-issues the request at most once, outside the 10-attempt budget. A second stall at that point ends the turn with
The response stalled before a response was produced. - The stream stopped mid-response. When the connection stays open but stops delivering data after Claude has completed a block of text or a tool call, the streaming idle watchdog (a timer that closes a connection that has gone quiet) aborts it, Claude Code keeps the completed output, and the notice reads
API Error: The response stopped arriving. The response above may be incomplete.
The watchdogs have their own variables. API_FORCE_IDLE_TIMEOUT controls the 5-minute body idle timeout that aborts a streaming response when no bytes arrive; when unset, that timeout is active on providers other than the direct Anthropic API, Claude Platform on AWS, and Amazon Bedrock with CLAUDE_ENABLE_BYTE_WATCHDOG_BEDROCK=1 set, and 0 turns it off for a slow gateway. CLAUDE_STREAM_IDLE_TIMEOUT_MS sets the event- and byte-level watchdog timeout, with a minimum of 300000 when you set it explicitly, and the byte-level watchdog caps it at 30 minutes. CLAUDE_ENABLE_STREAM_WATCHDOG is on by default for all providers when unset, since v2.1.196 [5].
The job's logs put a number on how often this happens for one job, read on 4 October 2026: across the 48 hourly runs from 10 to 12 September 2026, 7 failed, and all 7 were the API stopping its response part way through, which nothing in the job could fix and which the next hour's run retried on its own. Three days of one job is a record of one job, and the page on the job says so; the point here is that the failure Anthropic's reference describes is the one that appeared, and the next scheduled run was the retry.
Your own retry loop needs three numbers
Claude Code retries requests inside a turn. A retry of the whole run, when the process exits without a result, is your script's job, and it needs three numbers: an attempt count, a backoff between attempts, and a total budget. Reveneau's job, by its scripts read on 4 October 2026, uses AINEWS_WRITER_TRIES = 3, AINEWS_WRITER_BACKOFF_S = 60 seconds, and AINEWS_WRITER_BUDGET_S = 2,700 seconds, with each attempt capped at AINEWS_WRITER_TIMEOUT_S = 1,500 seconds, polled by the script because macOS ships no timeout command. A retry starts only if a full attempt plus its backoff still fits the remaining budget.
Why all three. A budget alone does not bound the attempt count: in an instant-failure test on 12 September 2026, the loop ran 14 attempts in 71 seconds before the count and the backoff were added. A count alone does not bound time: three attempts of 1,500 seconds plus two backoffs come to 4,620 seconds (our arithmetic: 3 times 1,500 plus 2 times 60), which overruns the hour. The fit rule ties them together: after one attempt that runs the full 1,500 seconds, 1,200 seconds of the 2,700 remain (our arithmetic), which is less than the 1,560 a further attempt and its backoff need, so no second capped attempt starts. The retries serve attempts that fail quickly, which is the case the error reference describes, and the budget stops a slow attempt from running into the next hour's run.
The job adds one more rule. A retry runs only against a clean working copy: if a failed attempt left files behind, the run stops instead of retrying with those files in place, because the retry would not know about a half-written item and the publish step would commit it beside the real work. The next hourly run resets the working copy anyway. Anthropic's documentation covers Claude Code's own retries of a request; the state of your files between attempts of your script is yours to check.
Each failure, what it costs, and the control
| Failure | What it costs, by the sources on this page | The control |
|---|---|---|
| Stalled stream before any text or tool call | One re-issued request at most, then the turn ends with an error [3] | The watchdog variables; CLAUDE_CODE_RETRY_WATCHDOG for a run that must wait out an outage [5] |
| No response headers within the deadline | One retry waiting one second less than API_TIMEOUT_MS, then No response from API [3] |
API_TIMEOUT_MS, CLAUDE_STREAM_FIRST_BYTE_TIMEOUT_MS [3] |
| Mid-response stall after a completed block | The completed output is kept; the turn continues from finished tool calls [3] | API_FORCE_IDLE_TIMEOUT, CLAUDE_STREAM_IDLE_TIMEOUT_MS [5] |
| Overloaded or throttled API (529, 429) | Up to 10 retries with exponential backoff; tokens up to the failure are consumed [3][9] | CLAUDE_CODE_MAX_RETRIES; CLAUDE_CODE_RETRY_WATCHDOG=1 in unattended runs [5] |
| Development server left running | Terminated after the final result, which Anthropic puts at five seconds later [1] | Nothing to set in Claude Code; your runner must still own any process holding its pipes (Reveneau's job) |
| Background command in an unattended session | Stopped at 30 minutes, or the passed timeout up to 2 hours, v2.1.285 or later [6] |
BASH_DEFAULT_TIMEOUT_MS, BASH_MAX_TIMEOUT_MS [6] |
| Subagent or workflow that never finishes | The run stays open; after 10 idle minutes Claude Code stops it and drops the partial result [1] | CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS; CLAUDE_ASYNC_AGENT_STALL_TIMEOUT_MS [5] |
| Run that keeps taking turns | No limit by default [2] | --max-turns, which exits with an error at the limit [2] |
| Whole run that never exits | Its scheduled hour, and every later run behind a lock (Reveneau's job) | A job-level timeout in CI [7][8]; an attempt cap in your own script |
| Script that retries a fast failure | Attempts limited only by the clock: 14 in 71 seconds in Reveneau's test | An attempt count and a backoff beside the budget (Reveneau's job) |
Our position
Give every unattended run two caps and one owner. The caps are --max-budget-usd for dollars and an outer time cap for the clock: a job timeout in CI, or an attempt cap your script enforces and polls. The owner is the process that started Claude Code, which ends the whole process group (the run and every process it started) when the run ends and writes the run's output to a file instead of a pipe. Inside the run, lower CLAUDE_CODE_MAX_RETRIES when the next scheduled run is the retry, and set CLAUDE_CODE_RETRY_WATCHDOG=1 only when there is no next run. Give your own retry loop an attempt count, a backoff and a budget, and let it retry only against a clean working copy. The page on running Claude Code from a script with a spend cap covers the dollar side, and the page on why usage rises in a long session covers what an interactive session left open costs, which is the same problem in an interactive session.
Reveneau is an AI software development consultancy. All of its code is written by AI, and every change must pass an eval suite, a set of automated tests written from the specification, before release, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. Every figure about Claude Code on this page is Anthropic's own statement about its own product, read on 4 October 2026, and every figure about the hourly job was read from that job's scripts and logs on the same day.
Common questions
Why does an unattended Claude Code run need a time cap as well as a spend cap?
Because `--max-budget-usd` counts dollars spent on API calls, and a run that is waiting spends none. Anthropic's CLI reference, read on 4 October 2026, describes the flag as the maximum dollar amount to spend before stopping; a stalled stream, an open background process or a subagent that never reports keeps the session open while that figure does not move. Reveneau's own hourly job, by its logs read the same day, once stayed open for 5 hours 47 minutes with its item finished and blocked five later runs.
What happens to a development server Claude started when a claude -p run ends?
Claude Code terminates it. Anthropic's non-interactive page, read on 4 October 2026, says a background Bash task started during a `claude -p` run, such as a dev server or a watch build, is terminated after Claude has returned its final result and standard input has closed, with a grace period Anthropic puts at five seconds so that a task finishing just after the result can still deliver its output. A background subagent or workflow is treated differently: the run stays open until that work completes or the idle wait ends.
How long does claude -p wait for a background subagent after the final turn?
By default the wait ends after 10 minutes of continuous idle waiting. Anthropic's non-interactive page, read on 4 October 2026, says `claude -p` stays open while a background subagent or workflow runs because its result is part of the final output, and that at the 10-minute point Claude Code stops whatever is still running and drops its partial result. `CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS` changes the limit, `0` waits without one, and the variable requires Claude Code v2.1.182 or later.
How many times does Claude Code retry a failed API request?
Up to 10 times, with exponential backoff, before it shows an error, by Anthropic's error reference read on 4 October 2026. The retried failures include server errors, overloaded responses and request timeouts that arrive before any of the response has streamed, dropped connections, and temporary 429 rate-limit responses. A failure that arrives after Claude has completed a block of text or a tool call is kept as partial output instead of retried, because re-running the request could execute the same tool calls twice.
What does CLAUDE_CODE_RETRY_WATCHDOG do?
Set to `1`, it makes Claude Code retry 429 and 529 capacity errors indefinitely instead of failing after `CLAUDE_CODE_MAX_RETRIES` attempts, with a wait of up to 5 minutes between attempts or until a reset time the response names. Anthropic's environment variable reference, read on 4 October 2026, says it is for unattended sessions such as CI jobs, that it fails at once on a 429 reporting a spend limit, and that on v2.1.199 or later it raises the retry count for other transient errors to 300. Requires v2.1.186 or later.
Should a script lower CLAUDE_CODE_MAX_RETRIES?
Anthropic's error reference, read on 4 October 2026, says to lower it to surface failures faster in scripts. The default is 10 retries and the value is capped at 15 as of v2.1.186. The choice depends on what the script does with a failure: a scheduler that runs the job again next hour, as Reveneau's own job does, can fail fast and let the next run retry, while a one-off CI job with no second chance may want `CLAUDE_CODE_RETRY_WATCHDOG=1` instead, which raises the count for transient errors.
What does The response stopped arriving mean in Claude Code?
It means the connection stayed open but stopped delivering data, so the streaming idle watchdog (a timer for quiet connections) aborted it, by Anthropic's error reference read on 4 October 2026. The message appears as `API Error: The response stopped arriving. The response above may be incomplete.` when the stall came after a completed block of text or tool call; Claude Code keeps what was completed, runs the finished tool calls, and continues the turn. A stall before any text or tool call is retried at most once.
What does API_TIMEOUT_MS control?
`API_TIMEOUT_MS` is the per-request timeout in milliseconds, 600000 by default, which is 10 minutes, by Anthropic's environment variable reference read on 4 October 2026. Anthropic says to raise it on slow networks or through a proxy, and its error reference adds that it also caps how long Claude Code waits for response headers: the retry after a `No response from API` failure waits one second less than `API_TIMEOUT_MS`. Values above 2147483647 overflow the timer and make requests fail at once.
What happens when I send SIGTERM to a claude -p run?
Claude Code exits with code 143 on SIGTERM, the standard stop signal, leaves the turn in progress unfinished and records no result for it, by Anthropic's non-interactive page read on 4 October 2026. On SIGTERM it terminates the process tree of any Bash command still running, runs `SessionEnd` hooks, and starts no new tool call or model request while exiting. To end the turn first, send SIGINT (what Ctrl+C sends) or call the Agent SDK's `interrupt()`. Set `CLAUDE_CODE_RESUME_INTERRUPTED_TURN=1` to have a later resume continue the interrupted turn.
Why does a time budget alone not bound the number of retry attempts?
Because an attempt that fails in seconds uses almost none of the budget, so the loop runs again at once. Reveneau's own hourly job, by its scripts and logs read on 4 October 2026, ran 14 attempts in 71 seconds in an instant-failure test on 12 September 2026 before an attempt count and a backoff were added. Its rule now is 3 attempts of 1,500 seconds each, 60 seconds apart, inside a 2,700-second budget, and a retry starts only if a full attempt plus its backoff still fits.
Why should a retry run only against a clean working tree?
Because a retry does not know what the failed attempt wrote. Reveneau's own hourly job, by its scripts read on 4 October 2026, stops instead of retrying when a failed attempt left files behind, since a second attempt would work with a half-written item in place and the publish step would commit both. The next hourly run resets the working copy and starts again. Anthropic's documentation covers Claude Code's own retries of API requests; the state of the files between attempts of your script is yours to check.
Is there a time limit on background commands in an unattended Claude Code session?
Yes. Anthropic's tools reference, read on 4 October 2026, says that in a session that runs unattended, such as a `-p` run, an Agent SDK application, a CI job or a cloud session, a background Bash or PowerShell command gets 30 minutes, or the `timeout` Claude passes with `run_in_background`, up to a maximum of 2 hours, counted from the moment it enters the background. Claude Code then stops it and tells Claude why. The limit requires Claude Code v2.1.285 or later; before v2.1.288 it applied in every session.
References
- Anthropic, Run Claude Code programmatically (code.claude.com), read 4 October 2026
- Anthropic, CLI reference (code.claude.com), read 4 October 2026
- Anthropic, Error reference (code.claude.com), read 4 October 2026
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Environment variables (code.claude.com), read 4 October 2026
- Anthropic, Tools reference (code.claude.com), read 4 October 2026
- Anthropic, Claude Code GitHub Actions (code.claude.com), read 4 October 2026
- Anthropic, Claude Code GitLab CI/CD (code.claude.com), read 4 October 2026
- Anthropic, Track cost and usage, Agent SDK (code.claude.com), read 4 October 2026
More in Scheduled and unattended
What a loop or a scheduled task costs in Claude Code while you are away
A loop or a scheduled task in Claude Code costs one full turn on every firing, and a turn sends the whole conversation to the model. A scheduled task is a prompt that Claude Code runs again by itself on a timer, with `/loop` or the cron tools. Anthropic's cost page, read on 4 October 2026, says a scheduled task fires on its interval even while the session is idle, sending your full context each time. The same page lists three more ways a session starts a turn with nobody typing: a goal check-in, a message from another of your sessions, and the short request behind a prompt suggestion. This page says how often each one fires, what Anthropic's documentation says it sends, and which setting stops it.
How to run Claude Code from a script with a spend cap
To run Claude Code from a script with a spend cap, call `claude -p` with your prompt and add `--max-budget-usd` with a dollar amount. `claude -p` is a run with no person at the keyboard: Claude Code reads the prompt, works, prints the result and exits. Anthropic's CLI reference, read on 4 October 2026, says `--max-budget-usd` is the maximum dollar amount to spend on API calls before stopping, works in print mode only, counts spend from subagents, and from Claude Code v2.1.217 fails any further subagent with `Budget limit reached` once the cap is reached. Add `--output-format json` so the run reports `total_cost_usd`, and `--model` with a full model name so the run stays on one model version. This page covers each flag, what it controls and its default.