The Claude Code prompt cache: how it works and what resets it
The Claude Code prompt cache is a store of request text that the service running the model has already processed. Claude Code sends the whole conversation with every request, and the cache lets the unchanged start of that request be billed as a cache read, which Anthropic's pricing page, read on 4 October 2026, lists at 0.1 times the base input price on most models. The match is exact, so a change early in the request recomputes everything after it. Anthropic's documentation lists nine actions that can reset the cache, such as switching models, and eight that keep it, such as editing files. A stored copy lasts five minutes or one hour without use. The /usage command shows how much of your input came from the cache.
Published October 4, 2026. Editorial.
Key takeaways
- Claude Code sends the full conversation with every request, and the prompt cache bills the unchanged start as a cache read, at 0.1 times the base input price on most models by Anthropic's pricing page read on 4 October 2026.
- Anthropic's documentation says the match is exact, so a change anywhere in the start of a request recomputes everything after it.
- Anthropic lists nine actions that can reset the cache, including switching models, turning on fast mode and compacting, and eight that keep it, including editing files and rewinding.
- A stored copy lasts five minutes or one hour without use, and Claude Code requests one hour by default only for the main conversation of a Claude subscription within plan usage.
- With Claude Code v2.1.251 or later, the Prompt cache (main) line in /usage shows the share of input tokens from cache, the misses and whether the cache is still within its lifetime.
Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. The work is counted in tokens, which are the pieces of text the model processes.
This guide is for a reader who saw one turn take longer than the turns before it, or saw a day of usage that was higher than expected. In both cases the first thing to understand is the prompt cache: what it stores, what makes the stored text unusable, and how to check it. Every fact here comes from Anthropic's documentation, read on 4 October 2026. The related guide on how to reduce Claude Code token usage covers what to change in your habits. This guide explains the mechanism behind several of those changes.
What the prompt cache is
The model keeps nothing between two requests. Anthropic's documentation says that Claude Code therefore sends the full context with every message: the system prompt, your project context, every earlier message and tool result, and your new message [1]. A tool result is what comes back when Claude reads a file or runs a command. New content is added at the end, so most of each request is identical to the one before it.
The prompt cache is a store, kept by the service that runs the model, of request text it has already processed. When a new request begins with the same text, the service reuses its earlier work. Text reused this way is billed as a cache read. Text stored for the first time is billed as a cache write.
Anthropic's pricing page, read on 4 October 2026, gives the prices as multiples of the base input price [3]:
- A cache read costs 0.1 times on most models, 0.05 times on Claude Opus 5.5, and 0.025 times on Claude Fable 5.1.
- A cache write costs 1.25 times when the stored copy is kept for five minutes, and 2 times when it is kept for one hour.
For Claude Sonnet 5.5, the API list price is $0.20 per million tokens for a cache read against $2 for base input [3]. Anthropic's pages give no statement of how cache tokens count against the usage limits of a subscription plan. Claude Code manages the cache for you. Anthropic's page says it "handles prompt caching for you, unless you disable it" [1]. The API reference adds that caching has no effect on the answer: the response is identical to the one you would get without it [2].
The page on what the prompt cache is covers the token types and the prices in full.
The exact-match rule
The cache works on the start of each request, which Anthropic calls the prefix. The service compares the prefix of a new request with text it processed recently. A request that finds a stored copy of its start is called a cache hit. A request that finds none is called a cache miss.
One sentence in Anthropic's documentation explains most of what follows: "The match is exact, so a change anywhere in the prefix recomputes everything after it." [1] The same page adds that there is no caching per file or per segment [1]. The store covers the start of the request as one piece.
To get the most from this rule, Claude Code puts the content that rarely changes first. Anthropic describes three layers [1]:
| Layer | What it contains | When it changes |
|---|---|---|
| System prompt | Anthropic's core instructions and the definitions of the tools the model can use | When the set of loaded tool definitions changes |
| Project context | CLAUDE.md (the file of instructions you write for Claude), auto memory (notes Claude saved for itself) and rules without a path limit | When a session starts, or after /clear or /compact |
| Conversation | Your messages, Claude's responses and tool results | On every turn |
A change in the conversation layer leaves the first two layers stored. A change in the system prompt makes everything after it unusable. To invalidate the cache means exactly this: to make the stored copy unusable for the next request.
Two settings are outside the table and still decide what stays stored. Each model has its own cache. And on most models, each effort level (the setting for how much the model reasons) has its own cache too [1]. The page on the three layers of a request explains each layer and the limits on sharing a stored copy between sessions.
The nine actions that reset the cache
Anthropic lists nine actions that can cause the next request to miss part or all of the cache. Its description of the result: "a one-time slower, more expensive turn, after which the new prefix is cached" [1]. The table is our summary of that list.
| Action | What the next request loses |
|---|---|
| Switching models | The whole request: it reads the entire conversation history with no cache hits |
| Changing effort level | The whole request on most models. By default the cache is kept on Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription |
| Turning on fast mode | The whole request, once per conversation |
| Connecting or removing an MCP server | Nothing while tool search holds back the MCP tools. The whole request when tools load in full at the start and a definition is added, or removed on purpose |
| Enabling or disabling a plugin | It depends on what the plugin provides. Its skills, commands, agents, hooks, monitors and themes keep the cache. Its MCP servers follow the row above |
| Denying an entire tool | Nothing while tool search is active. The whole request when tool search is unavailable or disabled |
| Compacting the conversation | The conversation layer, by design |
| Accumulating many images | The conversation, from the earliest message that held a removed image |
| Upgrading Claude Code | The first conversation after the upgrade builds its cache from the start |
Some terms in the table need one line each. Fast mode is a configuration that makes Claude Opus respond faster at a higher price per token [7]. An MCP server is a program that gives Claude extra tools, and tool search is the feature that holds back the definitions of those tools until Claude needs them. A plugin is a package of skills, hooks, subagents and MCP servers that you install as one unit, and a hook is a command that Claude Code runs automatically at a fixed point in its work. Compacting, or compaction, replaces the conversation history with a summary.
Claude Code warns you before some of these actions. When you run /model, it asks you to confirm the switch while the cache is still inside its lifetime and the new model differs from the one that produced the last response. Before v2.1.238 it asked even after the cache had expired. On most models it also asks before a change of effort level while the cache is warm, which means still inside its lifetime. And when a reload of plugins would cause a full re-read, it shows a warning and applies the reload only after you run /reload-plugins --force [1].
Our reading of these rules: a reset of the whole request costs more in a long conversation, because there is more text to process again. Anthropic makes this point about fast mode: turning it on at the start of a session costs less than turning it on late in a long one [1]. Anthropic's own sample shows the size. Its reference for hooks has a sample for a switch from Sonnet 5 to Opus 5 with 182,340 tokens of context. The sample's estimated cost of writing that context to the cache is $1.1396, and Anthropic says to treat it as an estimate [5]. By our arithmetic, that equals 182,340 tokens at the five-minute write price for Claude Opus 5, which the pricing page lists at $6.25 per million tokens [3]. The sample marks its price source as catalog, which the reference defines as list price [5].
Three pages cover the list. The nine actions that reset the cache is the full reference, with the version notes. Switching model or effort level in the middle of a task covers the first two rows and the switches that happen without the /model command. What /compact costs explains why compaction costs a fraction of the conversation size while the cache is in use, and the most after it has expired.
The eight actions that keep the cache
Anthropic also lists the actions that keep the cache. Each action on that list either adds text at the end of the conversation or leaves the request unchanged, so the stored start still matches [1].
| Action | Why the cache is kept |
|---|---|
| Editing files in your repository | File contents enter the request only when Claude reads them, and each read is added at the end |
| Editing CLAUDE.md during a session | The file was read once at session start. The edit applies after the next /clear, /compact or restart |
| Changing permission mode | The system prompt and the tool definitions stay the same |
| Changing output style | The new style's instructions arrive as a message in the conversation |
| Invoking skills and commands | Their instructions are added as user messages at the point where you invoke them |
Running /recap |
The summary is added as command output, and the history stays as it was |
| Rewinding the conversation | The history that remains is the same content the cache was built from |
| Spawning a subagent | The subagent's call and result are added at the end of the main conversation |
A permission mode sets which actions Claude can take without asking you. An output style changes the instructions that set the tone and format of Claude's replies. A skill is a file of instructions that Claude loads when it is relevant. A subagent is a helper that works on one task in its own conversation, with a cache of its own.
Three of these rows have a condition. With the opusplan model setting, entering or leaving plan mode (the mode in which Claude proposes a plan before it edits) switches between Opus and Sonnet, so it is a model switch. A skill or command whose settings name another model is a model switch for that turn. And an edit to CLAUDE.md has no effect on the running session until /clear, /compact or a restart [1]. The page on the eight actions that keep the cache explains each row.
How long a stored copy lasts
A stored copy is kept for a limited time without use. That time is called the time to live, or TTL. A cache inside its TTL is called warm, and one whose TTL has passed is called cold. Anthropic's documentation says each request that hits the cache starts the time again, so the cache stays warm for as long as you keep working. After a long enough gap, the next request recomputes the full input and stores it again [1].
There are two lifetimes: five minutes and one hour. Unless you choose one yourself, Claude Code requests one hour only for the main conversation of a Claude subscription that is within its plan usage. With usage credits, an API key or a cloud provider, the default is five minutes, and subagents get five minutes by default under both kinds of billing [1]. Usage credits are what a subscriber draws on after passing the plan's usage limit.
You can choose the lifetime with two settings, promptCacheTtl for the main conversation and subagentPromptCacheTtl for the other requests. Both require Claude Code v2.1.242 or later [1]. The one-hour lifetime costs more per write: 2 times the base input price against 1.25 times [3].
Resuming a session follows the same rules. Claude Code sends the whole conversation again, and the request reads from the cache whatever part of its start is unchanged and still inside the lifetime [1].
The page on five minutes or one hour has the default table, the order in which the controls apply, and an example of when the hour costs less.
How to check that the cache is working
Run /usage, the command that shows the token counts of the session. With Claude Code v2.1.251 or later, its Session block has a line named Prompt cache (main). Anthropic's page on costs, read on 4 October 2026, says the line gives the request count, the share of input tokens served from the cache, the misses, and whether the cache is warm [4]. Anthropic's sample line shows 14 requests, 91 percent of input tokens from cache, and 2 misses [4].
From v2.1.260 the line also names a likely cause for the last miss when Claude Code can identify one, for example likely cause: tool definitions changed [1]. Use it to connect a slow turn to one of the nine actions above.
Two more tools show the same numbers. The status line is a bar at the bottom of Claude Code that runs a script you choose, and the script can read the cache counts after every response. OpenTelemetry is an open standard for sending measurements to a monitoring system, and Claude Code's export reports cache read and cache creation tokens for each user and session [1].
The page on how to check your cache hit rate reads the /usage line field by field and includes a status line script.
Companies: gateways and cloud providers
The cache is kept wherever the model is served. With an API key or a Claude subscription, that is Anthropic's infrastructure. On Amazon Bedrock or Google Cloud's Agent Platform, it is the cloud provider's infrastructure [1].
A gateway is a proxy (a server that receives requests and passes them on) that an organisation runs between Claude Code and the model provider [6]. Through a gateway, the cache is wherever the requests are forwarded, and Anthropic says that whether caching works depends on the gateway [1]. Claude Code sends markers with each request that tell the service what to store. A gateway can forward them, reject the request, or remove them while returning success. In the last case, Anthropic's documentation says, the entire conversation history is billed as uncached input on every turn [1]. The gateway returns success in that case, so the request does not fail.
Two more limits matter to a company. The one-hour lifetime through a gateway depends on a header named anthropic-beta, a labelled line of information sent with the request, which the gateway has to forward unchanged. And on Amazon Bedrock, support for prompt caching and for the one-hour lifetime varies by model [1].
The page on gateways and cloud providers lists what to check for each connection and how to test a gateway.
Anthropic's advice, and ours
Anthropic puts its advice in one tip: "Pick your model and effort level at the top of a session, then save /compact for natural breaks between tasks. The fewer changes you make mid-task, the higher your cache hit rate." [1] The cache hit rate is the share of your input that is read from the store.
Our position adds three habits to that tip, each taken from the facts above.
- Set up before the first message. Choose the model and the effort level, turn on fast mode if you want it, connect the MCP servers you need, and enable the plugins you need. Those five rows of the table of nine actions are then settled before the task starts.
- Use the eight actions that keep the cache at any point in a task. Editing files, changing permission mode, invoking skills and rewinding keep the cache in the middle of a task. When an approach fails, rewind to an earlier turn: Anthropic's page says rewinding returns to a start that is already stored [1].
- Read the
/usageline after a long task. It counts the misses and, when Claude Code can identify one, names a likely cause for the last miss, so the next session can avoid it.
Reveneau is an AI software development consultancy. All of its code is written by AI, and every change must pass an eval suite (a set of automated tests) written from the specification before release, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. The figures in this guide are Anthropic's own statements about its own product, and the sums marked as ours are arithmetic on those figures.
Where to start
Start with the page that matches your question.
- You want the basic mechanism: what the prompt cache is, then the three layers of a request.
- One turn was slow: the nine actions that reset the cache and how to check your cache hit rate.
- You changed model or ran
/compact: switching model or effort level in the middle of a task and what /compact costs. - You want to know which actions keep the cache: the eight actions that keep the cache.
- You take breaks during a session: five minutes or one hour.
- Your company uses a gateway or a cloud provider: gateways and cloud providers.
- You want to use fewer tokens overall: how to reduce Claude Code token usage.
Explore the guide
How it works
What the Claude Code prompt cache is and why every message re-sends everything
The Claude Code prompt cache is a store, kept by the service that runs the model, of request text it has already processed. Claude Code needs it because the model keeps nothing between requests, so every message re-sends the whole conversation. When the start of a new request matches the stored text exactly, that part is billed as a cache read. On Anthropic's pricing page, read on 4 October 2026, a cache read costs 0.1 times the base input price on most models, and a five-minute cache write costs 1.25 times. Claude Code manages the cache automatically. This page explains the match rule, the token types and where the cache is stored.
The three layers of a Claude Code request and what changes each
A Claude Code request has three layers, in this order: the system prompt, the project context and the conversation. Anthropic's documentation, read on 4 October 2026, says Claude Code puts the content that rarely changes first, because the prompt cache matches the start of a request exactly. The system prompt changes when the set of loaded tool definitions changes. The project context changes when a session starts and after /clear or /compact. The conversation changes on every turn. A change in an early layer recomputes everything after it, so a change to the system prompt recomputes the whole request. This page explains each layer and where one stored copy can be reused.
What breaks it
The nine actions that reset the Claude Code prompt cache
Nine actions reset the Claude Code prompt cache, by Anthropic's documentation read on 4 October 2026: switching models, changing effort level, turning on fast mode, connecting or removing an MCP server, enabling or disabling a plugin, denying an entire tool, compacting the conversation, accumulating many images, and upgrading Claude Code. Each one can make the next request miss part or all of the stored text, which Anthropic describes as a one-time slower, more expensive turn. A model switch recomputes the whole request. An MCP server change and a tool deny rule keep the cache while tool search is on. This page lists what each action recomputes, the version notes, and how to avoid the cost.
What switching model or effort level in the middle of a task costs
Switching the model in the middle of a Claude Code task makes the next request read the entire conversation history with no cache hits, because each model has its own cache. That is Anthropic's documentation, read on 4 October 2026. Changing the effort level has the same result on most models, and keeps the cache on Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription. In Anthropic's own sample, a switch from Sonnet 5 to Opus 5 with 182,340 tokens of context has an estimated cache write cost of $1.1396. This page covers the confirmation prompt, the PreModelSwitch hook, the switches that happen without /model, and fast mode.
What /compact costs while the cache is active and after it has expired
/compact costs a fraction of what the conversation size suggests while the prompt cache is still active, and it costs the most after the cache has expired. Anthropic's documentation, read on 4 October 2026, explains why: to write the summary, Claude Code sends a separate request with the same system prompt, tools and history as your conversation, plus one instruction at the end. While the cache is still within its lifetime, that request reads the stored start from the cache. After a break longer than the cache lifetime, it reprocesses the full history as uncached input. In both cases the turn after compaction stores only the shorter history. This page compares the two cases and the alternatives.
Lifetime and checking
Five minutes or one hour: how long the Claude Code prompt cache lasts
The Claude Code prompt cache lasts five minutes or one hour without use, and each request that reads it starts the time again. Anthropic's documentation, read on 4 October 2026, says Claude Code asks for one hour only for the main conversation of a Claude subscription that is within its plan usage. With usage credits, an API key or a cloud provider, the default is five minutes. Subagents get five minutes by default. You can choose with the promptCacheTtl and subagentPromptCacheTtl settings, which require Claude Code v2.1.242 or later. On Anthropic's pricing page, a one-hour cache write costs 2 times the base input price and a five-minute write costs 1.25 times.
How to check your Claude Code cache hit rate
To check your Claude Code cache hit rate, run /usage and read the Prompt cache (main) line in the Session block. Anthropic's documentation, read on 4 October 2026, says the line shows the request count, the share of input tokens served from the cache, the number of misses, and whether the cache is warm, which means still inside its lifetime. It requires Claude Code v2.1.251 or later, and from v2.1.260 it also names a likely cause for the last miss when Claude Code can identify one. Anthropic's sample line shows 14 requests, 91 percent of input tokens from cache and 2 misses. A status line script can show the same numbers on every turn, and OpenTelemetry reports them for a whole organisation.
Common questions
What do warm and cold mean for the Claude Code prompt cache?
For the Claude Code prompt cache, warm means the stored copy is still inside its lifetime, and cold means the lifetime has passed. Anthropic's documentation, read on 4 October 2026, says each request that hits the cache starts the time again, so the cache stays warm for as long as you keep working. The lifetime is five minutes or one hour. After a long enough gap, the next request recomputes the full input and stores it again.
Where do the facts in this guide to the Claude Code prompt cache come from?
Every fact in this guide to the Claude Code prompt cache comes from Anthropic's own documentation, read on 4 October 2026: its page on how Claude Code uses prompt caching, its API reference, its pricing page, and its pages on costs, hooks, gateways and fast mode. The figures are Anthropic's statements about its own product. The sums marked as ours are arithmetic on those figures, and Reveneau is independent of Anthropic.
What does cache hit rate mean in Claude Code?
In Claude Code, the cache hit rate is the share of your input tokens that is read from the prompt cache. A request that finds a stored copy of its start is a cache hit, and one that finds none is a cache miss. Anthropic's sample of the `/usage` screen, read on 4 October 2026, shows 91 percent of input tokens from cache across 14 requests, with 2 misses. The line requires Claude Code v2.1.251 or later.
Does a cache reset cost more in a long conversation?
Yes. A reset of the whole request costs more in a long conversation, because there is more text to process again. Anthropic's documentation, read on 4 October 2026, says turning on fast mode at the start of a session costs less than turning it on late in a long one. In Anthropic's sample for a model switch with 182,340 tokens of context, the estimated cost of writing that context to the cache is $1.1396.
Does Claude Code warn me before an action that resets the cache?
Claude Code warns you before some actions that reset the cache. By Anthropic's documentation, read on 4 October 2026, it asks you to confirm a `/model` switch while the cache is still inside its lifetime, and on most models it asks before a change of effort level. When a reload of plugins would cause a full re-read, it shows a warning and waits for `/reload-plugins --force`. Before v2.1.238 the `/model` question appeared even after the cache had expired.
What does Anthropic recommend to keep the cache hit rate high?
Anthropic recommends choosing your model and effort level at the start of a session and keeping `/compact` for natural breaks between tasks. Its documentation, read on 4 October 2026, ends the tip with this sentence: "The fewer changes you make mid-task, the higher your cache hit rate." Reveneau adds one habit: connect MCP servers and enable plugins before the first message, so that those changes are made before the task starts.
Is the prompt cache different on a subscription and with an API key?
The prompt cache works the same way on a subscription and with an API key, and the default lifetime differs. Anthropic's documentation, read on 4 October 2026, says Claude Code requests the one-hour lifetime by default only for the main conversation of a Claude subscription within plan usage. With an API key, usage credits or a cloud provider, the default is five minutes. The `promptCacheTtl` setting changes it and requires Claude Code v2.1.242 or later.
Does resuming a session use the prompt cache?
Yes, for the part that still matches. Anthropic's documentation, read on 4 October 2026, says that when you resume a session, Claude Code sends the whole conversation again, and the request reads from the prompt cache whatever part of its start is unchanged and still inside the cache lifetime. After a gap longer than the lifetime, the next request recomputes the full input and stores it again.
How do I find out what reset my prompt cache?
To find out what reset your prompt cache, run `/usage` and read the `Prompt cache (main)` line. Anthropic's documentation, read on 4 October 2026, says the line shows the misses of the session, and from Claude Code v2.1.260 it names a likely cause for the last miss when Claude Code can identify one, for example `likely cause: tool definitions changed`. The line itself requires Claude Code v2.1.251 or later.
What should a company check about the prompt cache before it adopts Claude Code?
A company should check that its connection keeps the prompt cache working. Anthropic's documentation, read on 4 October 2026, says that through a gateway, whether caching works depends on the gateway. A gateway that removes the cache markers while returning success causes the entire conversation history to be billed as uncached input on every turn. On Amazon Bedrock, support for prompt caching and for the one-hour lifetime varies by model.
References
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Prompt caching (platform.claude.com), read 4 October 2026
- Anthropic, Pricing (platform.claude.com), read 4 October 2026
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Hooks reference (code.claude.com), read 4 October 2026
- Anthropic, Run Claude Code through a gateway (code.claude.com), read 4 October 2026
- Anthropic, Speed up responses with fast mode (code.claude.com), read 4 October 2026