Engineering

An unattended Claude Code job that stayed open for five hours and 47 minutes with nothing wrong with the model

Editorial · Reveneau · October 8, 2026

An unattended Claude Code job that stayed open for five hours and 47 minutes with nothing wrong with the model

An unattended run is a claude -p call started by a script, with no person at the keyboard. Claude Code is Anthropic's coding tool: you type a request in a terminal, the text window where you type commands, and an AI model reads files, runs commands and edits code. With -p, a script types the request instead, and Claude Code exits when the work is done. Everything the model reads and writes is counted in tokens, which are pieces of text.

Reveneau runs one such job every hour, and on 10 September 2026 one run stayed open for five hours and 47 minutes with its work finished after 14 minutes. Nothing was wrong with the model. This post tells that run from the job's own scripts and logs, read on 4 October 2026, then gives the changes that followed and the gap the job still has. One job is one job: none of the figures below is a statistic about Claude Code or about unattended runs.

What the job does

Reveneau is an AI software development consultancy, all of its code is written by AI, and every change must pass an eval suite, a set of automated tests written from the specification, before release. Its AI News section is written by an hourly job with nobody watching. The operating system's scheduler starts it at minute 7 of every hour. It writes a short news item, builds the site, commits and pushes, and the push deploys. No person reviews an item before it is published.

The steps, by the scripts read on 4 October 2026: the run takes a lock, so two runs never overlap. A plain script with no model checks the news sources and writes a candidate list, and zero candidates means the run exits without starting a model. claude -p runs once with a writing skill, a packaged set of instructions for one task, in acceptEdits permission mode, which lets Claude write files without asking, with --model opus pinned, with --output-format text, and with a fixed list of allowed tools. A second script writes a LinkedIn post for each new item, one claude -p call per item, with --output-format json. A build runs, and a publish script commits and pushes.

kill and pkill are deliberately absent from the writer's allowed tools. The next section is the reason.

The run that stayed open

The 17:07 PT run on 10 September 2026 stayed open for 5 hours 47 minutes and blocked the next five hourly runs. The writer had finished its item by 17:21. Its skill had asked it to confirm the new page rendered. No server was running, so it started a development server, which never exits. It tried to stop the server with pkill, which was absent from its allowed tools, so the call was refused. It wrote "That's OK; the process will get cleaned up" and stopped.

Nothing cleaned it up. The server had inherited the writer's output pipe, the channel that carried the writer's output to the runner. A pipe stays open until every process holding it has exited. So the runner waited on a server for five hours while the finished item stayed unpublished.

Anthropic's non-interactive documentation, read on 4 October 2026, says a background Bash task started during a claude -p run, such as a development server, is terminated after Claude has returned its final result and standard input has closed, with a grace period Anthropic puts at five seconds. The Claude Code version in use that day is outside what was read, so this post makes no claim about how that rule applied.

The four changes

  1. The writer's output goes to a file instead of a pipe. A leaked process can no longer hold the run open through the output channel.
  2. The runner ends the writer's whole process group after it exits, on every path including a clean one. The process group is the writer and every process it started.
  3. A per-attempt time cap, polled by the script because the operating system has no timeout command. The value came from measurement: across 40 healthy writer passes from 10 to 12 September 2026, the shortest took 323 seconds, the median 648, the 90th percentile 952 and the longest 1,201. None reached 1,500 seconds, so 1,500 became the cap, down from 1,800.
  4. The runner owns the render server. It starts the server used to check the page before the writer runs, tells the writer where it is, and stops it afterwards. The writer is told never to start a server.

The next run after the change: render server ready in 2 seconds, writer 7 minutes 38 seconds, 256 of 256 candidates evaluated, build OK, 10 minutes 5 seconds in total.

A second round followed on 12 September 2026. Across the 48 hourly runs from 10 to 12 September, 41 were healthy and 7 failed, and all 7 failures were the API stopping its response part way through. The writer now retries inside the run: up to 3 attempts, 60 seconds apart, inside a total budget of 2,700 seconds, and a retry starts only if a full attempt plus its wait still fits. The attempt count and the wait matter beside the budget, because an instant-failure test that day ran 14 attempts in 71 seconds before they were added. The page on timeouts, retries and cleanup for an unattended run gives Anthropic's own time controls at each level.

What the job records about cost

The writer step is called with --output-format text. Anthropic's documentation describes text as the default plain text output, and says that with --output-format json the response includes total_cost_usd and a per-model cost breakdown. The job therefore records no token count and no dollar figure for its main step. This is its largest gap, and the fix is one flag.

The LinkedIn step uses json and records total_cost_usd per call. Measured on 21 September 2026 over 6 calls: $0.78 to $2.33 per item, in 29 to 256 seconds. These are figures Claude Code reports at list price for the opus model, and Anthropic's Agent SDK page calls total_cost_usd a client-side estimate that can differ from the bill.

No spend cap is set on either call. Anthropic's CLI reference describes --max-budget-usd as the maximum dollar amount to spend on API calls before stopping, in print mode only, with subagent spend counted, with enforcement from Claude Code v2.1.217. A subagent is a second copy of Claude working in its own context window, and a writing skill is free to start one. The page on running Claude Code from a script with a spend cap covers the flag.

Reveneau recommends two changes to its own job, in this order: change the writer's call to --output-format json and record total_cost_usd from every run, then add --max-budget-usd to both calls with a figure above the largest healthy run. For the writer no figure exists yet, so a cap set today would be a guess.

A run that waits spends nothing and still costs the hour

A spend cap counts dollars, and a run that is waiting spends none. A stalled response, a server that never exits, or a subagent that never reports keeps a run open while the dollar counter stands still. That is why this job's caps are time caps, and why it needs the dollar cap as well. The full account is on the page about an hourly unattended job and its measured cost, and the guide to the cost of agents and unattended runs in Claude Code gives the controls from Anthropic's documentation. Reveneau is independent of Anthropic, whose documentation, read on 4 October 2026, supplied every statement here about what the flags do.

A run that costs nothing while it waits still costs the hour.

Sources

Common questions

What happened in Reveneau's unattended Claude Code job on 10 September 2026?

By the job's logs read on 4 October 2026, the 17:07 PT run stayed open for 5 hours 47 minutes and blocked the next five hourly runs. The writer had finished its item by 17:21. Its skill asked it to confirm the new page rendered, no server was running, so it started a development server, which never exits. The server inherited the writer's output pipe, and the runner waited until every process holding the pipe had gone.

Why did the writer leave a development server running?

Because it could not stop it. The writer tried pkill, which was absent from its allowed tool list, so the call was refused, and it wrote "That's OK; the process will get cleaned up" and stopped, by the logs read on 4 October 2026. Anthropic's documentation says a claude -p run with no permission host denies a request that would otherwise prompt, which is what that refused call was. The runner has owned the render server since.

Does Anthropic's documentation say what happens to a background process when claude -p ends?

Anthropic's non-interactive page, read on 4 October 2026, says a background Bash task started during a claude -p run, such as a development server, is terminated after Claude has returned its final result and standard input has closed, with a grace period Anthropic puts at five seconds. The Claude Code version in use on 10 September 2026 is outside what was read for the write-up, so it makes no claim about how that rule applied that day.

What four changes followed the hang?

The writer's output goes to a file instead of a pipe. The runner ends the writer's whole process group, every process it started, after it exits, on every path including a clean exit. A per-attempt time cap, polled by the script, limits each writer attempt. And the runner starts and stops the server used to check the page itself, and tells the writer never to start a server. The next run after the change took 10 minutes 5 seconds in total.

How was the 1,500-second cap chosen?

From measurement. Across 40 healthy writer passes from 10 to 12 September 2026, by the logs read on 4 October 2026, the shortest took 323 seconds, the median 648, the 90th percentile 952 and the longest 1,201. None reached 1,500 seconds, so an attempt still running at that point is stuck, and the cap came down from 1,800. Three attempts are allowed, 60 seconds apart, inside a total budget of 2,700 seconds.

Why does a retry loop need an attempt count and a wait, beside a time budget?

Because a time budget alone bounds nothing when each attempt fails in seconds. In an instant-failure test on 12 September 2026, by the logs read on 4 October 2026, the loop ran 14 attempts in 71 seconds before an attempt count and a backoff were added. The rule now is 3 attempts of up to 1,500 seconds, 60 seconds apart, inside 2,700 seconds, and a retry starts only if a full attempt plus its wait still fits the remaining budget.

Why is the cost of the job's main step unmeasured?

Because the writer is called with --output-format text, which Anthropic's documentation, read on 4 October 2026, describes as the default plain text output, so the job records no token count and no dollar figure for it. Anthropic says that with --output-format json the response includes total_cost_usd and a per-model cost breakdown. The job's second step already uses json and recorded $0.78 to $2.33 per call over 6 calls on 21 September 2026.

Does the job set a spend cap with --max-budget-usd?

No. Neither claude -p call in the job sets --max-budget-usd, by its scripts read on 4 October 2026; every cap in the job is a time cap. Anthropic's CLI reference describes the flag as the maximum dollar amount to spend on API calls before stopping, print mode only, with subagent spend counted and enforcement from Claude Code v2.1.217. Reveneau recommends adding it to both calls once the writer's cost has been measured, with the figure set above the largest healthy run.

Why does the write-up say one job is one job?

Because every figure comes from one job's scripts and logs, read on 4 October 2026, on the days named, with one skill. Forty writer passes over three days and six calls on one day describe that job then, and nothing about another job or another week. The figures are given so a reader can see what measuring an unattended run looks like, including the gap this one left, and none of them is a statistic about Claude Code or about unattended runs.