What breaks it

The nine actions that reset the Claude Code prompt cache

Nine actions reset the Claude Code prompt cache, by Anthropic's documentation read on 4 October 2026: switching models, changing effort level, turning on fast mode, connecting or removing an MCP server, enabling or disabling a plugin, denying an entire tool, compacting the conversation, accumulating many images, and upgrading Claude Code. Each one can make the next request miss part or all of the stored text, which Anthropic describes as a one-time slower, more expensive turn. A model switch recomputes the whole request. An MCP server change and a tool deny rule keep the cache while tool search is on. This page lists what each action recomputes, the version notes, and how to avoid the cost.

Published October 4, 2026. Editorial.

Key takeaways

  • Anthropic's documentation, read on 4 October 2026, lists nine actions that can cause the next Claude Code request to miss part or all of the prompt cache.
  • Switching models means the next request reads the entire conversation history with no cache hits, because each model has its own cache.
  • Turning on fast mode costs one cache miss per conversation, and turning it off and on again later keeps the cache, by Anthropic's documentation.
  • With tool search holding back the MCP tools, a server that connects or disconnects during a session does not disturb anything already cached.
  • Claude Code applies an upgrade at the next launch and never in the middle of a session, so the cost is an uncached first turn after a restart.

Claude Code is Anthropic's coding tool: you type a request in a terminal, and an AI model reads files, runs commands and edits code. Each message becomes a new request that contains the whole conversation again, counted in tokens, the pieces of text that the model processes.

The prompt cache is a store of request text that the service has already processed. When the start of a new request matches the stored text exactly, the service reads that part from the store at a lower price. Anthropic's name for the start of a request is the prefix. A cache hit is a request that finds its start in the store, and a cache miss is a request that finds none of it, or only part of it.

To invalidate the cache means to make the stored copy unusable for the next request. This page calls that a reset. Anthropic's documentation, read on 4 October 2026, lists nine actions that "can cause the next request to miss part or all of the cache" [1]. This page covers all nine. It is the reference list in the guide to the Claude Code prompt cache.

What a reset costs

A reset costs you once. Anthropic describes it this way: "You see a one-time slower, more expensive turn, after which the new prefix is cached." [1] A turn is one message from you plus the work Claude does in reply.

The size of that cost depends on where the change is. Claude Code builds each request in three parts: the system prompt (Anthropic's instructions and the tool definitions), the project context (your instruction files) and the conversation. The match is exact, so a change in an early part recomputes everything after it [1]. The page on the three layers of a request explains the order.

Anthropic adds: "Most of them are avoidable mid-task once you know they have a cost." [1]

The nine actions in one table

Each row is from Anthropic's documentation, read on 4 October 2026 [1]. The last column is our summary of how to avoid the cost, taken from the same page.

Action What is recomputed How to avoid it
1. Switching models The whole request. The next request reads the entire conversation history with no cache hits Choose the model before the first message
2. Changing effort level The whole request on most models. By default the cache is kept on Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription. The section below lists where that exception stops Choose the level before the first message
3. Turning on fast mode The whole request, once per conversation Turn it on at the start of the session
4. Connecting or removing an MCP server Nothing when tool search defers the MCP tools. The cache is invalidated when tools load in full at the start and a definition is added, or removed on purpose Keep tool search on, and set up servers before the session
5. Enabling or disabling a plugin It depends on what the plugin provides. Skills, commands, agents, hooks, monitors and themes keep the cache. MCP servers follow row 4 Read the warning that /reload-plugins shows before a full re-read
6. Denying an entire tool Nothing when tool search is active. The cache is invalidated when tool search is unavailable or disabled Use a scoped rule where one is enough
7. Compacting the conversation The conversation part, by design Run /compact at a natural break between tasks
8. Accumulating many images The conversation, from the earliest message that held a removed image Send fewer or smaller images
9. Upgrading Claude Code The first conversation after the upgrade builds its cache from the start An upgrade applies at the next launch. Set DISABLE_AUTOUPDATER=1 to choose when

1 and 2. Switching models and changing effort level

Each model has its own cache. Anthropic's page says that switching with /model, the command that changes the model, "means the next request reads the entire conversation history with no cache hits, even though the content is identical" [1]. Three other events count as a model switch on the same page: entering or leaving plan mode (the mode in which Claude proposes a plan before it edits) with the opusplan model setting, which uses Opus in plan mode and Sonnet for execution; an automatic model fallback on Fable models, Opus 5.5, Sonnet 5.5 and Opus 5, in which Claude Code runs a flagged request again on another model; and a skill or command that names a different model [1].

When you run /model in the terminal, Claude Code asks you to confirm the switch only while the cache is still within its lifetime and the new model is different from the one that produced the last response. The cache counts as within its lifetime for one lifetime period after Claude Code last sent a request in the conversation or Claude last responded. Before v2.1.238, Claude Code did not check the lifetime and asked even after the cache had expired [1].

The effort level is the setting for how much the model reasons before it replies. On most models, changing it in the middle of a session has the same result as a model switch, and Claude Code asks you to confirm the change while the cache is still within its lifetime. On Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription, changing effort keeps the cache, and Claude Code applies the new level without asking. That exception stops in five places: on Amazon Bedrock, on Google Cloud's Agent Platform, on a Claude apps gateway (Anthropic's own gateway, a server that passes your requests on to a provider), when you set the environment variable CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS, and when your organisation has a HIPAA configuration. Before v2.1.260, changing effort on Fable 5.1 with an API key or a Claude subscription also invalidated the cache [1].

The page on switching model or effort level in the middle of a task covers both topics in full and gives a worked example.

3. Turning on fast mode

Fast mode is a setting that makes Claude Opus respond faster at a higher price per token. Anthropic's fast mode page, read on 4 October 2026, says it is supported on Opus 5.5, Opus 5 and Opus 4.8 [3].

Turning it on adds a header to the request. A header is a labelled value sent with the request, outside the conversation text. Anthropic says this header "is part of the cache key", which means it is one of the values the service uses to identify a stored copy. The first request that Claude Code sends with fast mode on therefore "reads the entire conversation history with no cache hits" [1].

Four details follow, all from the same two pages [1] [3]:

  • Timing. Claude Code sets the header once when a turn starts and keeps it for the whole turn. If you turn fast mode on while Claude is working, the cache miss happens on the first request of your next turn.
  • Price. The uncached input tokens of that request are billed at fast mode rates. That is Anthropic's stated reason why turning fast mode on at the start of a session costs less than turning it on late in a long one.
  • A second cost. If your current model does not support fast mode, turning it on also switches your model, and that switch starts a new cache from the next request in the running turn.
  • Once per conversation. After the first fast mode turn, Claude Code keeps sending the header and changes only the speed setting of the request, which is outside the cache key. Turning fast mode off, the automatic return to standard speed after a rate limit, and turning it on again later all keep the cache. /clear and /compact reset this state, because they rebuild the cache at those points.

4. Connecting or removing an MCP server

An MCP server is a program that connects Claude Code to an outside tool or data source. Its tools arrive as tool definitions, the descriptions that tell the model what each tool does. Anthropic's page says the definitions are in the system prompt part, "so the cache invalidates when the set of tool definitions in the request changes between turns" [1].

Whether a server change does that depends on tool search. Tool search is a feature that holds back the full definitions of MCP tools until Claude needs them. Anthropic's MCP page, read on 4 October 2026, says it is enabled by default [2].

  • With the tools held back, Claude Code keeps the tool list from the conversation's first request for the whole conversation. A server that connects or disconnects during the session does not disturb anything already stored [1].
  • With the tools loaded in full at the start, adding a definition invalidates the cache, and so does removing one on purpose [1].

The second case applies when tool search is below its auto threshold, disabled or unavailable. Anthropic gives three examples of unavailable: Google Cloud's Agent Platform models earlier than the Claude 4.5 generation, a gateway reached through a custom ANTHROPIC_BASE_URL, and a Microsoft Foundry deployment hosted on Azure once Claude Code detects that the deployment rejects tool search [1]. A gateway is a server that receives your requests first and passes them on.

For sessions without tool search, Anthropic gives a table of changes in the middle of a session [1]:

Change in the middle of a session Cache
A server connects, or a server adds tools while connected Invalidated
A server stops with no action on your part, such as a local server's process exiting Kept
A remote server reconnects automatically after its connection is lost Kept, with one exception: a request sent during the reconnection can add a tool named WaitForMcpServers, which invalidates the cache once
You remove a tool on purpose, with a deny rule or by disabling its server in /mcp Invalidated

Two more facts from the same page. Editing your MCP configuration file leaves the cache as it is, because the new configuration takes effect only after a restart. Turning the advisor tool on or off with /advisor also keeps the cache, because its definition is placed after the point where the stored start ends [1].

5. Enabling or disabling a plugin

A plugin is a package that adds features to Claude Code. What a plugin change costs "depends on which component types the plugin provides" [1].

  • Skills, commands, agents, hooks, monitors and themes. A skill is a file of instructions that Claude loads when it is relevant, and a hook is a command that Claude Code runs automatically at a fixed point in its work. Anthropic's wording is that Claude Code "never invalidates the cache" for these. It adds their content after the existing conversation, so the next request pays for that content and still reads everything before it from the cache [1].
  • MCP servers. A plugin that provides MCP servers follows the rules of action 4 [1].
  • Code intelligence plugins. Each of these connects Claude Code to a language server for one programming language, so that Claude can see type errors and find definitions in the code [5]. Anthropic's paragraph on this case says one thing: when you enable such a plugin, Claude gets the LSP tool [1].

A change made in the /plugin menu is applied by the /reload-plugins command, which Claude Code runs for you when you close the menu. You pay the cost on the first turn after the change applies. When the reload would cause a full re-read, Claude Code shows a warning and does not apply the reload. Running /reload-plugins --force applies it anyway [1].

The page gives three version notes [1]:

  • Moving the session to another folder with /cd on v2.1.246 or later applies the plugins that the new folder's settings enable, without the full re-read warning.
  • Adding or removing a plugin in a folder passed with --plugin-dir applies immediately in interactive sessions. If it would cause a full re-read, Claude Code keeps the change waiting and shows a notice to run /reload-plugins. This requires Claude Code v2.1.265 or later.
  • /reload-plugins also runs in sessions without an interactive terminal when you type it in directly. This requires Claude Code v2.1.260 or later. In those sessions, plugin MCP server changes take effect in your next session.

The change can be undone. When you disable a plugin that you enabled earlier in the same session, Claude Code restores the earlier form of the request. If that stored start is still within its lifetime, the next request reads the older cache entry [1].

6. Denying an entire tool

A deny rule is a permission rule that blocks Claude from using a tool. If you add a bare tool name such as Bash or WebFetch as a deny rule, Claude cannot call that tool from your next request on. This holds whether you add the rule through the /permissions command or by editing a settings file [1].

The effect on the cache again depends on tool search. When tool search is active, the tool definitions in the request stay the same and the stored start is kept. When tool search is unavailable or disabled, Claude Code removes the definition from the next request, which invalidates the cache. Removing the rule later invalidates it again [1].

Only a rule that matches the whole tool does this: a bare tool name, the equivalent form Bash(*), or a pattern that matches tool names. A pattern that matches only MCP tools, such as mcp__*, blocks those tools in the same way. Scoped deny rules such as Bash(rm *), and all allow and ask rules, leave the list of tools unchanged. Claude Code checks them when Claude tries a call, which leaves the stored start intact [1].

7. Compacting the conversation

Compaction replaces your message history with a summary. Anthropic's page says: "By design, this invalidates the conversation layer, since the next request has a new, shorter history that doesn't share a prefix with the old one." [1] Claude Code reuses the system prompt part, with one exception for a resumed conversation that the page linked below explains. It reloads the project context from disk, and that reloaded text is read from the cache only if your instruction files and memory are unchanged since the session started [1].

The page on what /compact costs covers this topic in full.

8. Accumulating many images

The API, the service that receives each request, limits how many images and PDF files one request can contain. Claude Code also limits their total size, so large screenshots reach the limit with fewer images than small ones [1].

When the next request would pass either limit, Claude Code removes a group of the oldest images and PDF files from what it sends. Claude can no longer see the removed images. Removing them changes the messages that contained them, so the next request reprocesses the conversation from the earliest of those messages onward [1]. The API reference states the general rule: adding or removing images anywhere in the request affects the cached messages [4].

Because Claude Code removes a group at a time, Anthropic says, you see one slower turn per group instead of one with each new screenshot [1].

9. Upgrading Claude Code

A new version of Claude Code, in Anthropic's words, "typically updates the system prompt or tool definitions". The first conversation you start after an upgrade therefore builds its cache from the start [1].

The timing helps you. Automatic updates download in the background and apply at the next launch, "never mid-session" [1]. You see the cost as an uncached first turn after a restart. To choose when upgrades apply, set the environment variable DISABLE_AUTOUPDATER=1 [1].

A conversation that you started before the upgrade and then resume is a separate case. By default it keeps the system prompt it started with, and the new prompt takes effect once the conversation is compacted or in a new conversation [1].

Which resets to avoid in the middle of a task

Our position: plan for three of the nine before the task starts, and accept the other six when the work needs them.

Decide the model, the effort level and fast mode before the first message. A model switch and the first use of fast mode recompute the whole request, and so does an effort change on most models. Each of the three is your own choice. Anthropic's tip on the same page says to choose your model and effort level at the start of a session, and its fast mode page says to enable fast mode at the start of a session for the best cost efficiency [1] [3].

Keep tool search on. With it, actions 4 and 6 keep the cache. The related guide's page on MCP servers and command-line tools covers the setting.

Compact at a break. Compaction resets the conversation part by design, and Anthropic's advice is to run it between tasks [1].

When a turn is slow and you do not know why, read the /usage screen. It has a line named Prompt cache (main) that counts the misses in the session. The line requires Claude Code v2.1.251 or later, and on v2.1.260 or later it also names a likely cause for the last miss when Claude Code can identify one [1]. The page on how to check your cache hit rate reads that line field by field.

Reveneau is an AI software development consultancy. All of its code is written by AI, so token use is a running cost of every Reveneau build, and a reset in the middle of a long task is part of that cost. Reveneau is independent of Anthropic. Every statement about the product on this page is Anthropic's own, read on 4 October 2026.

Common questions

What resets the Claude Code prompt cache?

Nine actions reset the Claude Code prompt cache, in whole or in part. Anthropic's documentation, read on 4 October 2026, lists them: switching models, changing effort level, turning on fast mode, connecting or removing an MCP server, enabling or disabling a plugin, denying an entire tool, compacting the conversation, accumulating many images, and upgrading Claude Code. Anthropic says most of them are avoidable in the middle of a task once you know they have a cost.

What happens on the turn after the prompt cache is reset?

The turn after the prompt cache is reset is slower and costs more, and after it the new start of the request is stored. Anthropic's documentation, read on 4 October 2026, describes a reset as a one-time slower, more expensive turn. How much is recomputed depends on where the change is: a model switch recomputes the whole request, and compaction invalidates the conversation part by design.

Does a crashed MCP server reset the prompt cache?

No. An MCP server that stops with no action on your part keeps the prompt cache. Anthropic's documentation, read on 4 October 2026, gives the example of a local server's process exiting and marks the cache as kept in sessions without tool search. With tool search holding back the MCP tools, Claude Code keeps the tool list from the conversation's first request for the whole conversation, so a server that disconnects does not disturb anything already stored.

Does editing my MCP configuration file reset the prompt cache?

No. Editing your MCP configuration file leaves the prompt cache as it is. Anthropic's documentation, read on 4 October 2026, says the new configuration takes effect only after a restart. A server that connects in a running session is a different event: in a session where the tools load in full at the start, an added tool definition invalidates the cache, and so does a definition you remove on purpose.

Does enabling a plugin reset the prompt cache?

Enabling a plugin resets the prompt cache only for some kinds of plugin content. Anthropic's documentation, read on 4 October 2026, says Claude Code never invalidates the cache for a plugin's skills, commands, agents, hooks, monitors or themes, because it adds their content after the existing conversation. A plugin that provides MCP servers follows the same rules as connecting or removing an MCP server.

What does /reload-plugins --force do?

`/reload-plugins --force` applies a plugin reload that Claude Code has declined to apply because of its effect on the prompt cache. Anthropic's documentation, read on 4 October 2026, says that when a reload would cause a full re-read of the conversation, Claude Code shows a warning and does not apply the reload. Running the command with `--force` applies it anyway, and you pay the cost on the first turn after the change applies.

Does a deny rule for one command reset the prompt cache?

No. A scoped deny rule such as `Bash(rm *)` keeps the prompt cache. Anthropic's documentation, read on 4 October 2026, says scoped deny rules and all allow and ask rules leave the list of tools unchanged, and Claude Code checks them when Claude tries a call. A rule that blocks a whole tool, such as the bare name `Bash`, invalidates the cache when tool search is unavailable or disabled.

Why did a turn become slow after I shared many screenshots?

A turn becomes slow after many screenshots because Claude Code removed a group of the oldest images from the request, and that changed earlier messages. Anthropic's documentation, read on 4 October 2026, says the next request then reprocesses the conversation from the earliest of those messages onward. Claude Code removes a group at a time, so you see one slower turn per group. Claude can no longer see the removed images.

Does a Claude Code update reset the cache in the middle of a session?

No. A Claude Code update applies at the next launch and never in the middle of a session. Anthropic's documentation, read on 4 October 2026, says automatic updates download in the background, and the first conversation you start after the upgrade builds its cache from the start. You see that as an uncached first turn after a restart. Setting `DISABLE_AUTOUPDATER=1` lets you choose when upgrades apply.

Does turning the advisor tool on or off reset the prompt cache?

No. Turning the advisor tool on or off with `/advisor` keeps the prompt cache. Anthropic's documentation, read on 4 October 2026, gives the position of the advisor tool's definition as the reason: it is placed after the point where the stored start of the request ends. Other tool definitions are in the system prompt part, where a change in the set of definitions between turns invalidates the cache.

Does turning fast mode off and on again reset the cache each time?

No. Fast mode costs one cache miss per conversation, on the first request sent with fast mode on. Anthropic's documentation, read on 4 October 2026, says that after the first fast mode turn Claude Code keeps sending the header and changes only the speed setting, which is outside the cache key. Turning fast mode off, the automatic return to standard speed after a rate limit, and turning it on again later all keep the cache.

Which cache resets can I plan before a task starts?

You can plan three cache resets before a task starts: the model, the effort level and fast mode. Anthropic's tip, read on 4 October 2026, says to choose your model and effort level at the start of a session, and its fast mode page says to enable fast mode at the start of a session for the best cost efficiency. Reveneau also recommends keeping tool search on, because with it an MCP server change and a tool deny rule keep the cache.

More in What breaks it