The Claude Code prompt cache: how it works and what resets it / Lifetime and checking
Five minutes or one hour: how long the Claude Code prompt cache lasts
The Claude Code prompt cache lasts five minutes or one hour without use, and each request that reads it starts the time again. Anthropic's documentation, read on 4 October 2026, says Claude Code asks for one hour only for the main conversation of a Claude subscription that is within its plan usage. With usage credits, an API key or a cloud provider, the default is five minutes. Subagents get five minutes by default. You can choose with the promptCacheTtl and subagentPromptCacheTtl settings, which require Claude Code v2.1.242 or later. On Anthropic's pricing page, a one-hour cache write costs 2 times the base input price and a five-minute write costs 1.25 times.
Published October 4, 2026. Editorial.
Key takeaways
- Anthropic's documentation, read on 4 October 2026, says each request that reads the prompt cache starts its lifetime again, so the cache stays usable for as long as you keep working.
- By default Claude Code requests the one-hour lifetime only for the main conversation of a Claude subscription within plan usage, and five minutes with usage credits, an API key or a cloud provider.
- The promptCacheTtl and subagentPromptCacheTtl settings accept 5m or 1h and require Claude Code v2.1.242 or later.
- On Anthropic's pricing page, read on 4 October 2026, a one-hour cache write costs 2 times the base input price and a five-minute cache write costs 1.25 times.
- Running claude -p "hello" --output-format json reports one-hour cache writes under ephemeral_1h_input_tokens and five-minute cache writes under ephemeral_5m_input_tokens.
Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. The work is counted in tokens, which are the pieces of text the model processes. The model keeps nothing between requests, so Claude Code sends the whole conversation again with every request.
The prompt cache is a store, kept by the service that runs the model, of request text it has already processed. When a new request begins with exactly the same text, the service reuses its earlier work and bills that text at a lower price. A stored copy is kept only for a limited time. This page explains how long, who decides, and what the longer option costs. It is part of the guide to the Claude Code prompt cache.
What the lifetime is and when it starts again
The time to live, or TTL, is how long an unused stored copy is kept. Anthropic's documentation, read on 4 October 2026, says the API (the service that receives each request) offers two: a five-minute TTL and a one-hour TTL [1]. A cache that is still inside its lifetime is called warm. A cache whose lifetime has passed is called cold.
The lifetime counts time without use. The Claude Code page says: "Each request that hits the cache resets the timer, so the cache stays warm as long as you keep working." [1] A request hits the cache when it finds a stored copy of its start. The API reference adds that this refresh has no extra charge [2].
After a long enough gap, the same page says, "the next request recomputes the full input and re-establishes the cache" [1]. That is the slow first turn after a break.
One detail in the API reference changes how you count. The lifetime is measured from the start of the request that writes or reads the stored copy, and the time Claude spends writing its answer is included. Anthropic gives an example: if a response takes 4 minutes to arrive, the next request that reuses the same stored text has to start within what Anthropic puts at 1 minute of that response finishing [2]. The example is for the five-minute lifetime.
Two groups of requests
Claude Code decides the lifetime for each request. Anthropic's documentation, read on 4 October 2026, says every request belongs to one of two fixed groups, which it calls buckets [1]:
- The main conversation. This is your turns in an interactive session, runs started with
claude -p(the mode that executes one prompt and exits), and turns sent through the Agent SDK, Anthropic's packages for running the same agent from a program. It also includes the helper requests that Claude Code runs as part of those turns. - Everything else. These are the requests Claude Code makes outside that conversation. Anthropic's examples are subagents (helpers that work on one task in their own conversation), workflows, in-process teammates, forks, compaction (the step that replaces the conversation history with a summary) and session titles.
The two groups have different defaults, and each has its own control.
The default lifetime depends on how you are billed
Unless you choose a lifetime yourself, Claude Code asks for one hour only in one case: a Claude subscription that is still within the usage included in the plan. There it asks for the hour for the main conversation, and for a small set of helper requests that Anthropic controls on its servers [1]. The table is Anthropic's, from its documentation read on 4 October 2026 [1].
| Group of requests | Claude subscription, within plan usage | Usage credits, API key or cloud provider |
|---|---|---|
| Main conversation | One hour | Five minutes |
| Everything else | Five minutes, except the helper requests that Anthropic controls on its servers, which get one hour | Five minutes |
An API key is the credential of an account that pays per token. A cloud provider here means a service such as Amazon Bedrock that runs Claude models for you.
Usage credits change the default. Usage credits let you keep working after you pass your plan's usage limit [9]. Anthropic's page says that once Claude Code draws on usage credits, "you are billed for that usage, so Claude Code drops the main conversation to the cheaper five-minute TTL" [1]. A subscriber who passes the limit can therefore find that a ten-minute break now ends with a cold cache. The ten minutes are our invented example, and the rule is Anthropic's. To keep the one-hour lifetime after that point, Anthropic says to choose the lifetime yourself, with one of the controls in the next section [1].
Subagents are the second case to note in the table. They are outside the main conversation group, so they get five minutes even on a subscription, until you choose a longer lifetime for them [1].
How to choose the lifetime yourself
You can set a lifetime for either group. Each control accepts 5m or 1h, and Claude Code ignores any other value [1]. There are two kinds of control. A setting is a key in a Claude Code settings file such as settings.json. An environment variable is a named value that a program reads from the shell when it starts.
| Group of requests | Setting | Environment variable |
|---|---|---|
| Main conversation | promptCacheTtl |
CLAUDE_CODE_PROMPT_CACHE_TTL |
| Everything else | subagentPromptCacheTtl |
CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL |
Both settings and both environment variables require Claude Code v2.1.242 or later [1]. Anthropic's settings reference, read on 4 October 2026, gives this example, which keeps the main conversation on one hour and subagents on five minutes [4]:
{
"promptCacheTtl": "1h",
"subagentPromptCacheTtl": "5m"
}
The settings reference says both keys can be set in any settings file, and that both are unset by default, so each request gets the default from the table above [4]. It also says that subagentPromptCacheTtl covers subagents, workflows, and Claude Code's own background and helper requests, such as compaction and session titles [4].
For a reader who signs in with an API key or uses a cloud provider, Anthropic's instruction is direct: set promptCacheTtl to 1h to give the main conversation a one-hour cache. Requests outside it stay on five minutes until you choose a lifetime for that group too [1].
Which control applies when several are set
When more than one control applies to a request, Claude Code takes the first match in this order. The list is Anthropic's, read on 4 October 2026 [1]:
FORCE_PROMPT_CACHING_5M=1, which forces five minutes for both groups.- The group's environment variable.
- The group's setting.
- For a subagent's requests, the
cacheTtlvalue in theexperimentalfield of the subagent's frontmatter, the block of settings at the top of its file. This requires Claude Code v2.1.248 or later, and Claude Code ignores a1hthere while your Claude subscription is using usage credits [1] [6]. ENABLE_PROMPT_CACHING_1H=1, which requests one hour for both groups.- The default for the request's group.
Two of these variables act on both groups at once.
FORCE_PROMPT_CACHING_5M is first in the order. Anthropic names three uses for it: debugging cache behaviour, comparing the two lifetimes, and overriding a longer lifetime that was set in managed settings [1]. Managed settings are settings that an organisation's administrators enforce for all of its users.
ENABLE_PROMPT_CACHING_1H is fifth in the order, so any control above it in the order is applied first. Anthropic's list of environment variables says it is intended for people who use an API key, Amazon Bedrock, Google Cloud's Agent Platform, Microsoft Foundry or Claude Platform on AWS. Subscribers who are drawing on usage credits can set it to keep the one-hour lifetime [5]. The same list marks an older variable, ENABLE_PROMPT_CACHING_1H_BEDROCK, as deprecated [5].
An organisation can set one policy for everyone by putting these variables in the env block of managed settings [1].
What the one-hour lifetime costs
The one-hour lifetime costs more per write. Text stored for the first time is billed as a cache write, and a one-hour write costs more than a five-minute write. Stored text used again is billed as a cache read. The multipliers come from Anthropic's pricing page, read on 4 October 2026 [3]:
| Operation | Price relative to base input |
|---|---|
| Five-minute cache write | 1.25 times |
| One-hour cache write | 2 times |
| Cache read | 0.1 times on most models, 0.05 times on Claude Opus 5.5, 0.025 times on Claude Fable 5.1 and Claude Mythos 5.1 |
In dollars per million tokens, the same page lists Claude Sonnet 5.5 at $2 for base input, $2.50 for a five-minute write, $4 for a one-hour write and $0.20 for a cache read. It lists Claude Opus 5.5 at $4, $5, $8 and $0.20 [3].
The pricing page also states when each lifetime recovers its extra write cost. Where a cache read costs 10 percent of the input price, the extra cost of a five-minute write is recovered after one cache read, and the extra cost of a one-hour write after two [3].
Here is an invented example, with our own arithmetic on the Claude Sonnet 5.5 list prices. Suppose the stored start of a conversation is 100,000 tokens. The example ignores the new text of each turn and the model's output.
- Writing 100,000 tokens costs $0.25 at the five-minute price and $0.40 at the one-hour price.
- Reading 100,000 tokens from the cache costs $0.02.
Now compare two cases. In the first, you send the next message two minutes later. With either lifetime the next request is a cache read. The five-minute path costs $0.25 plus $0.02, which is $0.27. The one-hour path costs $0.40 plus $0.02, which is $0.42. The one-hour path costs $0.15 more, and the extra lifetime is not used.
In the second case, you send the next message after a 20-minute break. On the one-hour path the request is still a cache read, and the total is again $0.42. On the five-minute path the stored copy has expired. Anthropic's page says the request "recomputes the full input and re-establishes the cache", and it gives no price for that request [1]. If we assume that all 100,000 tokens are written again at the five-minute price, the total is $0.25 plus $0.25, which is $0.50.
These dollar figures are API list prices. Anthropic's pages give no statement of how cache writes count against the usage limits of a subscription plan.
When the hour is worth choosing
Anthropic's own summary, read on 4 October 2026, matches the example. The longer lifetime helps "when you leave a session idle and come back to it". It costs more on short periods of work that never pause for longer than five minutes, where the higher write price applies and the extra lifetime goes unused [1].
The API reference gives the same rule in different words. If the same text is used more often than every 5 minutes, keep the five-minute cache, because each use refreshes it at no extra charge. The one-hour cache is for text that is used less often than every 5 minutes and more often than every hour. One of Anthropic's examples is a side agent that takes longer than 5 minutes [2].
Our position follows from that. If you pay per token and your working day includes pauses of between five minutes and an hour (reviews, meetings, waiting for a test run), set promptCacheTtl to 1h. If you work in short uninterrupted sessions, stay on the default. Leave subagentPromptCacheTtl unset unless you can name a subagent or workflow whose own requests are more than five minutes apart.
A pause longer than an hour ends with a cold cache under either lifetime. The related guide's page on why usage rises in a long session covers what to do before a long break, and the page on what /compact costs covers resuming an old session.
How to confirm which lifetime you got
Anthropic gives one command for this, in its documentation read on 4 October 2026 [1]:
claude -p "hello" --output-format json
Read usage.cache_creation in the result. Claude Code reports one-hour cache writes under ephemeral_1h_input_tokens and five-minute cache writes under ephemeral_5m_input_tokens [1]. The field with a number above zero shows the lifetime that your main conversation's cache writes used.
Inside a session, the /usage screen shows the same fact in words. Anthropic's sample of its Prompt cache (main) line ends with "warm (1h TTL, last activity 40s ago)" [9]. The page on how to check your cache hit rate reads that line field by field.
Three places where the hour is limited
The one-hour lifetime is offered on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry, by Anthropic's API reference read on 4 October 2026 [2]. Three conditions apply in Claude Code:
- A gateway you reach through
ANTHROPIC_BASE_URL. A gateway is a server that your organisation runs between Claude Code and the model provider. Part of the one-hour request is sent in a header namedanthropic-beta. A header is a labelled line of information sent with a request. The gateway has to forward that header unchanged [1]. - The Claude apps gateway. This is Anthropic's own gateway product. The one-hour lifetime is not available through it, and prompt caching through it uses the five-minute lifetime [7].
- Amazon Bedrock. Support for prompt caching, the minimum size of a stored start, and the availability of the one-hour lifetime all vary by model [1]. Anthropic's Bedrock page shows
ENABLE_PROMPT_CACHING_1H=1as the optional way to request the hour there [8].
The page on gateways and cloud providers covers these cases.
Our position in one paragraph
Decide the lifetime from how you pause. Reveneau recommends the five-minute default for work without breaks and promptCacheTtl set to 1h for per-token accounts whose sessions include breaks of five minutes to an hour, because that is the range Anthropic names for the one-hour cache [2]. Then run the claude -p check once, so that you know which lifetime your requests received.
Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. Every figure on this page is Anthropic's own statement about its own product, and the sums marked as ours are arithmetic on Anthropic's list prices.
Common questions
What does TTL mean for the Claude Code prompt cache?
TTL means time to live: how long an unused stored copy in the Claude Code prompt cache is kept. Anthropic's documentation, read on 4 October 2026, says the API offers a five-minute TTL and a one-hour TTL. A cache inside its lifetime is called warm, and one whose lifetime has passed is called cold. After the lifetime passes, the next request recomputes the full input and stores it again.
Does using the prompt cache restart its lifetime?
Yes. Each request that reads the prompt cache starts its lifetime again. Anthropic's documentation, read on 4 October 2026, says the cache stays usable for as long as you keep working, and the API reference says this refresh has no extra charge. The lifetime is measured from the start of a request, so the time Claude spends writing an answer is included in it.
Which Claude Code requests get the one-hour cache lifetime by default?
By default, only the main conversation of a Claude subscription within its plan usage gets the one-hour cache lifetime, together with a small set of helper requests that Anthropic controls on its servers. That is Anthropic's documentation, read on 4 October 2026. Subagents, workflows, compaction and session titles get five minutes. With usage credits, an API key or a cloud provider, both groups of requests get five minutes.
Why did my cache lifetime change from one hour to five minutes?
Your cache lifetime changes from one hour to five minutes when a Claude subscription passes its plan's usage limit and Claude Code starts to draw on usage credits. Anthropic's documentation, read on 4 October 2026, says you are billed for that usage, so Claude Code moves the main conversation to the cheaper five-minute lifetime. To keep one hour, set `promptCacheTtl` to `1h` or set `ENABLE_PROMPT_CACHING_1H=1`.
How do I get a one-hour cache lifetime with an API key?
To get a one-hour cache lifetime with an API key, set `promptCacheTtl` to `1h` in a Claude Code settings file. Anthropic's documentation, read on 4 October 2026, gives this instruction for people who sign in with an API key or use a cloud provider. The setting requires Claude Code v2.1.242 or later. It covers the main conversation only, so requests outside it stay on five minutes until you also set `subagentPromptCacheTtl`.
What does subagentPromptCacheTtl cover?
The `subagentPromptCacheTtl` setting covers the requests Claude Code makes outside the main conversation. Anthropic's settings reference, read on 4 October 2026, lists subagents, workflows, and Claude Code's own background and helper requests, such as compaction and session titles. The setting accepts `5m` or `1h`, is unset by default and requires Claude Code v2.1.242 or later. Its environment variable is `CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL`.
Which control applies when several cache lifetime controls are set?
When several cache lifetime controls are set, Claude Code takes the first match in a fixed order. By Anthropic's documentation, read on 4 October 2026, the order is `FORCE_PROMPT_CACHING_5M=1`, then the group's environment variable, then the group's setting, then a subagent's `cacheTtl` frontmatter value, then `ENABLE_PROMPT_CACHING_1H=1`, then the default for the request's group. The frontmatter value requires Claude Code v2.1.248 or later.
What does FORCE_PROMPT_CACHING_5M do?
Setting `FORCE_PROMPT_CACHING_5M=1` forces the five-minute cache lifetime for both groups of requests, and it is first in Claude Code's order of controls. Anthropic's documentation, read on 4 October 2026, names three uses: debugging cache behaviour, comparing the two lifetimes, and overriding a longer lifetime set in managed settings. Managed settings are settings that an organisation's administrators enforce for all of its users.
What does ENABLE_PROMPT_CACHING_1H do?
Setting `ENABLE_PROMPT_CACHING_1H=1` requests the one-hour cache lifetime for both groups of requests. It is fifth in Claude Code's order of controls, so a group's environment variable or setting is applied before it. Anthropic's list of environment variables, read on 4 October 2026, says it is intended for people who use an API key, Amazon Bedrock, Google Cloud's Agent Platform, Microsoft Foundry or Claude Platform on AWS.
How much more does a one-hour cache write cost than a five-minute cache write?
A one-hour cache write costs 2 times the base input price, and a five-minute cache write costs 1.25 times, on Anthropic's pricing page read on 4 October 2026. For Claude Sonnet 5.5 the page lists $4 per million tokens for a one-hour write and $2.50 for a five-minute write. In our invented example of 100,000 stored tokens, that is $0.40 against $0.25.
When is the one-hour cache lifetime worth the higher write price?
The one-hour cache lifetime is worth its higher write price when you leave a session idle for more than five minutes and come back within the hour. Anthropic's documentation, read on 4 October 2026, says the longer lifetime costs more on short periods of work that never pause past five minutes. The API reference says to keep the five-minute cache for text used more often than every 5 minutes.
How do I confirm which cache lifetime my session used?
To confirm which cache lifetime your main conversation used, run `claude -p "hello" --output-format json` and read `usage.cache_creation` in the result. Anthropic's documentation, read on 4 October 2026, says Claude Code reports one-hour cache writes under `ephemeral_1h_input_tokens` and five-minute cache writes under `ephemeral_5m_input_tokens`. Inside a session, the `Prompt cache (main)` line of `/usage` also names the lifetime in effect.
References
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Prompt caching (platform.claude.com), read 4 October 2026
- Anthropic, Pricing (platform.claude.com), read 4 October 2026
- Anthropic, All settings (code.claude.com), read 4 October 2026
- Anthropic, Environment variables (code.claude.com), read 4 October 2026
- Anthropic, Create custom subagents (code.claude.com), read 4 October 2026
- Anthropic, Claude apps gateway (code.claude.com), read 4 October 2026
- Anthropic, Claude Code on Amazon Bedrock (code.claude.com), read 4 October 2026
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026