The Claude Code prompt cache: how it works and what resets it / What keeps it
Eight actions that keep the Claude Code prompt cache
Eight actions keep the Claude Code prompt cache, by Anthropic's documentation read on 4 October 2026: editing files in your repository, editing CLAUDE.md during a session, changing permission mode, changing output style, invoking skills and commands, running /recap, rewinding the conversation, and spawning a subagent. Each one either adds text at the end of the conversation or leaves the request unchanged, so the stored start of the request still matches. Three of them have a condition attached. A CLAUDE.md edit does not apply until /clear, /compact or a restart. With the opusplan setting, a change into or out of plan mode is a model switch. A skill that names another model is a model switch for that turn.
Published October 4, 2026. Editorial.
Key takeaways
- Anthropic's documentation, read on 4 October 2026, lists eight actions that keep the Claude Code prompt cache, because each one adds text at the end of the conversation or leaves the request unchanged.
- Project-root and user-level CLAUDE.md files are read once at session start, so an edit keeps the cache and does not apply until the next /clear, /compact or restart.
- Changing permission mode keeps the cache, except with the opusplan model setting, where entering or leaving plan mode switches between Opus and Sonnet.
- Before Claude Code v2.1.251, a change of output style during a session kept the cache and did not apply until /clear or a new session.
- A subagent builds its own cache with a five-minute lifetime by default, and a fork reads the main session's cache on its first request.
Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. The work is counted in tokens, which are the pieces of text the model processes. The model keeps nothing between requests, so Claude Code sends the whole conversation again with every request.
The prompt cache is a store, kept by the service that runs the model, of request text it has already processed. When a new request begins with exactly the same text, the service reuses its earlier work and bills that text at a lower price. Anthropic's documentation, read on 4 October 2026, lists eight actions that keep this store usable [1]. This page explains each one: why the stored text is kept, and what to check. It is part of the guide to the Claude Code prompt cache. The opposite list is on the page about the nine actions that reset the cache.
The one rule behind all eight actions
The cache works on the start of each request. Anthropic's name for that start is the prefix. The service compares the prefix of a new request with text it processed recently. A request that finds a stored copy of its start is called a cache hit, and a request that finds none is called a cache miss. To invalidate the cache means to make the stored copy unusable for the next request.
The documentation states the rule in one sentence: "The match is exact, so a change anywhere in the prefix recomputes everything after it." [1]
Claude Code builds each request in a fixed order. The system prompt comes first: Anthropic's instructions for the model and the definitions of the tools the model can use. The project context comes second: text loaded from your project, such as CLAUDE.md, the file of instructions that you write for Claude. The conversation comes last [1]. The page on the three layers of a request describes each part.
Anthropic gives one reason for the whole list. Each of them either adds text at the end of the conversation or leaves the request unchanged [1]. Text added at the end leaves everything before it identical, so the stored start still matches. The next request reads the earlier text from the cache and processes only the new text.
The eight actions in one table
The actions and the reasons are Anthropic's, from its page on prompt caching [1]. The third column collects the conditions that Anthropic attaches to each one.
| Action | Why the cache is kept | What to check |
|---|---|---|
| Editing files in your repository | File contents enter the request only when Claude reads them, and each read is added at the end of the conversation | The earlier read stays in the conversation as it was. Claude Code adds a note that the file changed |
| Editing CLAUDE.md during a session | The project-root and user-level files are read once at session start and held in memory | The edit does not apply until the next /clear, /compact or restart |
| Changing permission mode | The system prompt and the tool definitions stay the same | With the opusplan model setting, entering or leaving plan mode is a model switch |
| Changing output style | The new style's instructions arrive as a message in the conversation | Before v2.1.251 the new style did not apply until /clear or a new session |
| Invoking skills and commands | Their instructions are added as user messages at the point where you invoke them | A skill or command that names a different model can be a model switch for that turn |
Running /recap |
The summary is added as command output, and the message history stays as it was | Claude Code limits the output to 400 characters |
| Rewinding the conversation | The history that remains is the same content the cache was built from at that point | The two Summarize options in the rewind menu are a different operation |
| Spawning a subagent | The subagent's call and its result are added at the end of the main conversation | The subagent builds a separate cache of its own, with a five-minute lifetime by default |
Editing files in your repository
You can edit any file in your project while a session runs, and the cache stays intact. Anthropic's documentation, read on 4 October 2026, gives the reason: "File contents enter context only when Claude reads them, and reads append to the conversation." [1] Context here means the text the model reads in one request.
A file that Claude already read is a special case. The text of that earlier read is part of the conversation history, and your edit leaves it there unchanged. Claude Code adds a system reminder that says the file changed, and Claude reads the file again if it needs to [1]. A system reminder is a message that Claude Code adds to the conversation to give Claude information. Anthropic's glossary lists "a note that a file Claude read earlier has changed on disk" among the messages that reach Claude this way [7].
Both the note and the second read are new text at the end of the conversation. The stored start still matches.
Editing CLAUDE.md during a session
Editing CLAUDE.md keeps the cache, and the reason is that the edit has no effect on the running session. Anthropic's documentation, read on 4 October 2026, says the project-root and user-level CLAUDE.md files "are read once at session start and held in memory", and that Claude keeps working with the version that was loaded at session start. The new content loads on the next /clear, /compact or restart [1]. /clear starts a new conversation. /compact replaces the conversation history with a summary, a step that Anthropic calls compaction.
Two kinds of instruction file follow a different schedule. A nested CLAUDE.md is one that is kept in a subdirectory of your project. A path rule is a rules file whose paths: field limits it to matching files. Anthropic's page on memory says a nested file is included when Claude reads files in that subdirectory, and a path rule loads when Claude uses the Read, Write or Edit tool on a matching file [2]. The prompt caching page adds what that means for edits. Editing one of these files before it loads does take effect. After it loads, its content is part of the conversation history, and a later edit leaves that history as it was [1].
There is a delayed cost to know about. When you run /compact, Claude Code loads the project context from disk again. Anthropic's page says that this reload is read from the cache only if CLAUDE.md and memory are unchanged since the session started [1]. Our reading: an edited CLAUDE.md therefore changes the project context part of the request at the moment the edit begins to apply.
Our advice: edit CLAUDE.md between tasks, then run /clear. The related guide's page on how long CLAUDE.md should be covers what belongs in the file.
Changing permission mode or output style
A permission mode sets which actions Claude can take in a session without asking you first [4]. You change it with Shift+Tab in the terminal, and you can change it at any time during a session [4]. Anthropic's documentation, read on 4 October 2026, says that switching between modes "does not change the system prompt or tool definitions", so a mode change keeps the cache [1].
Plan mode is the permission mode in which Claude explores the code and proposes a plan before it changes anything [4]. Plan mode adds its instructions as conversation messages, so entering it also keeps the stored start [1].
There is one exception. The opusplan model setting uses the Opus model during plan mode and the Sonnet model during execution. Each model has its own cache, so with opusplan every change into or out of plan mode is a model switch [1]. The page on switching model or effort level in the middle of a task covers what a model switch costs.
An output style is a configuration that changes the instructions Claude Code gives Claude, to set the behaviour, tone or format of its replies [7]. You can switch styles during a session with /output-style, with /config or with the outputStyle setting. Claude uses the new style from your next message. Claude Code delivers the new style's instructions as a message in the conversation, so that request still reads the system prompt and the earlier conversation from the cache [1].
Anthropic records a version change here. Before v2.1.251, a style switch during a session kept the cache, and the new style did not apply until you ran /clear or started a new session [1].
Skills, commands and /recap
A skill is a file of instructions or a workflow that Claude loads when it is relevant, or that you invoke by name [7]. A command is an instruction you type with a leading slash. Anthropic's documentation, read on 4 October 2026, says both "inject their instructions as user messages at the point of invocation", and that nothing earlier in the conversation changes [1]. The cache is kept, and the request processes the new instructions as new text.
Check one field. Skills and commands read their settings from frontmatter, a block of settings at the top of the file [7]. When the frontmatter names a model other than the session's current model, that turn is a model switch: the next request reads the entire conversation history with no cache hits. The session model returns on your next prompt [1].
/recap writes a short summary of the session for you to read in the terminal. Anthropic's page says it adds the summary as command output and keeps your message history, so the stored start stays intact [1]. The page on interactive mode says Claude Code limits recap output to 400 characters [6]. /compact is the opposite case: it replaces the history, and the page on what /compact costs covers it.
Rewinding the conversation
/rewind takes the conversation back to an earlier turn. You open the menu with /rewind, or by pressing Esc twice when the input is empty [5]. Anthropic's documentation, read on 4 October 2026, explains why the cache is kept: "The remaining history is the same content the cache was built from at that point, and the system prompt and project context layers are unchanged, so the next request hits the earlier cache entry." [1]
A stored copy expires when nobody uses it. The time an unused copy is kept is called the time to live, or TTL. A cache that is still inside that time is called warm. Anthropic's page says that rewinding works even when the turn you return to is older than the TTL: "Every turn since then has read through that prefix, which kept the entry warm." [1]
The rewind menu can also restore files to their state at that turn. Anthropic's page says that restoring files alongside the conversation has no separate effect on the cache, because file contents enter the request only when Claude reads them [1].
The same menu has two Summarize options, and Anthropic's page on checkpointing compares them to a targeted /compact [5]. The statement about keeping the cache covers going back to an earlier turn. Anthropic's prompt caching page gives no statement of what the Summarize options do to the cache.
Starting a subagent
A subagent is a helper that works on one task in its own conversation and returns a summary to the main conversation. Anthropic's documentation, read on 4 October 2026, says a subagent "starts its own conversation with its own system prompt and tool set, separate from the parent's" [1]. Two things follow.
The main conversation keeps its cache. From its side, the subagent's call and result are added at the end of the conversation [1].
The subagent reads nothing from the main conversation's cache on its first request, because the two requests start with different text. It builds a cache of its own over its turns. By default that cache has a five-minute lifetime, including on a Claude subscription [1]. The page on how long the cache lasts explains how to choose a longer one.
A fork is the exception. A fork is a subagent that inherits the whole conversation so far. It has the same system prompt, tools, model and message history as the main session, so its first request reads the main session's cache [1] [3]. Anthropic's page on subagents says this makes a fork cheaper than a new subagent for tasks that need the same context [3]. You start a fork yourself with /subtask, which requires Claude Code v2.1.212 or later [3].
Anthropic lists three more requests that can read text an earlier request stored [1]:
- A session copy. When you copy a session with
/fork, the copy receives its extra instruction as a message at the end of the copied conversation, so the cache that the original conversation built stays intact. - A resumed subagent. When Claude resumes a subagent, the first request of the resumed run can read the cache that the original run built.
- A workflow that starts several agents with the same start. Claude Code makes all but the first agent wait for up to 5 seconds by default, so that their first requests can read the text the first agent stored.
Our position
Use these eight actions at any point in a task, and check the three conditions that turn one of them into a model switch or a delayed change. First, with opusplan, treat Shift+Tab into plan mode as a model switch. Second, read the frontmatter of a skill before you use it in a long session, and look for a model field. Third, expect a CLAUDE.md edit to apply only after /clear, /compact or a restart.
When you want a helper that already knows the conversation, Reveneau recommends a fork over a new subagent, for the reason Anthropic gives: the fork's first request reads the stored text of the main session [3].
Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. Every statement about Claude Code on this page is Anthropic's own description of its product, read on 4 October 2026.
Common questions
Which actions keep the Claude Code prompt cache?
Eight actions keep the Claude Code prompt cache, by Anthropic's documentation read on 4 October 2026. They are editing files in your repository, editing CLAUDE.md during a session, changing permission mode, changing output style, invoking skills and commands, running `/recap`, rewinding the conversation, and spawning a subagent. Each one adds text at the end of the conversation or leaves the request unchanged, so the stored start of the request still matches.
Why do some actions keep the prompt cache while others reset it?
An action keeps the prompt cache when it leaves the start of the request identical. Anthropic's documentation, read on 4 October 2026, says the match is exact, so a change anywhere in the start recomputes everything after it. The eight actions add text at the end of the conversation or leave the request unchanged. A model switch is the opposite case: each model has its own cache, so the next request reads the whole history with no cache hits.
Does editing a file that Claude already read reset the prompt cache?
No. Editing a file that Claude already read keeps the prompt cache. Anthropic's documentation, read on 4 October 2026, says file contents enter the request only when Claude reads them, and each read is added at the end of the conversation. The text of the earlier read stays in the history as it was. Claude Code adds a system reminder that the file changed, and Claude reads the file again if it needs to.
Does Claude see my edit to a file if the earlier read stays in the conversation?
Yes. Claude learns about your edit from a note that Claude Code adds to the conversation. Anthropic's documentation, read on 4 October 2026, says Claude Code adds a system reminder that the file changed, and Claude reads the file again if needed. A system reminder is a message Claude Code adds to give Claude information. The note and the new read are both added at the end, so the prompt cache is kept.
Does editing CLAUDE.md in the middle of a session reset the prompt cache?
No. Editing CLAUDE.md in the middle of a session keeps the prompt cache, and the edit has no effect on the running session. Anthropic's documentation, read on 4 October 2026, says project-root and user-level CLAUDE.md files are read once at session start and held in memory. The new content loads on the next `/clear`, `/compact` or restart. After `/compact`, the project context is read from the cache only if CLAUDE.md and memory are unchanged.
Does an edit to a nested CLAUDE.md file apply during the session?
An edit to a nested CLAUDE.md file applies during the session only if you make it before the file loads. Anthropic's documentation, read on 4 October 2026, says a nested file loads when Claude reads files in that subdirectory, and a rule with a `paths:` field loads when Claude first reads a matching file. After the file loads, its content is part of the conversation history, and a later edit leaves that history as it was.
Does changing the permission mode reset the prompt cache?
No. Changing the permission mode keeps the prompt cache, because the system prompt and the tool definitions stay the same. That is Anthropic's documentation, read on 4 October 2026. The exception is the `opusplan` model setting, which uses Opus during plan mode and Sonnet during execution. With `opusplan`, entering or leaving plan mode is a model switch, and each model has its own cache.
Does changing the output style reset the prompt cache?
No. Changing the output style during a session keeps the prompt cache. Anthropic's documentation, read on 4 October 2026, says Claude Code delivers the new style's instructions as a message in the conversation, so the request still reads the system prompt and the earlier conversation from the cache. Claude uses the new style from your next message. Before v2.1.251, the new style did not apply until you ran `/clear` or started a new session.
Does rewinding still read from the cache when the earlier turn is older than the cache lifetime?
Yes. Rewinding reads from the cache even when the turn you return to is older than the cache lifetime. Anthropic's documentation, read on 4 October 2026, gives the reason: every turn since then has read through that same start of the request, which kept the stored copy in use. The history that remains after `/rewind` is the same content the cache was built from at that point, so the next request hits the earlier cache entry.
Does starting a subagent reset the main conversation's prompt cache?
No. Starting a subagent keeps the main conversation's prompt cache. Anthropic's documentation, read on 4 October 2026, says that from the main conversation's side, the subagent's call and its result are added at the end of the conversation. The subagent has its own system prompt and tool set, so it builds a separate cache over its turns. That separate cache has a five-minute lifetime by default, including on a Claude subscription.
What is the difference between a fork and other subagents for the prompt cache?
A fork reads the main session's prompt cache on its first request, and other subagents start a separate cache. Anthropic's documentation, read on 4 October 2026, says a fork inherits the same system prompt, tools, model and message history as the main session. Anthropic says this makes a fork cheaper than a new subagent for tasks that need the same context. You start a fork with `/subtask`, which requires Claude Code v2.1.212 or later.
Does copying a session with /fork keep the prompt cache?
Yes. Copying a session with `/fork` keeps the prompt cache that the original conversation built. Anthropic's documentation, read on 4 October 2026, says the copy receives its extra instruction as a message at the end of the copied conversation, so the stored start stays intact. The same page lists two more cases that read stored text: a resumed subagent, and a workflow whose later agents wait up to 5 seconds by default.
References
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, How Claude remembers your project (code.claude.com), read 4 October 2026
- Anthropic, Create custom subagents (code.claude.com), read 4 October 2026
- Anthropic, Choose a permission mode (code.claude.com), read 4 October 2026
- Anthropic, Checkpointing (code.claude.com), read 4 October 2026
- Anthropic, Interactive mode (code.claude.com), read 4 October 2026
- Anthropic, Glossary (code.claude.com), read 4 October 2026