The Claude Code prompt cache: how it works and what resets it / What breaks it
What switching model or effort level in the middle of a task costs
Switching the model in the middle of a Claude Code task makes the next request read the entire conversation history with no cache hits, because each model has its own cache. That is Anthropic's documentation, read on 4 October 2026. Changing the effort level has the same result on most models, and keeps the cache on Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription. In Anthropic's own sample, a switch from Sonnet 5 to Opus 5 with 182,340 tokens of context has an estimated cache write cost of $1.1396. This page covers the confirmation prompt, the PreModelSwitch hook, the switches that happen without /model, and fast mode.
Published October 4, 2026. Editorial.
Key takeaways
- Each model has its own cache, so after a /model switch the next request reads the entire conversation history with no cache hits, by Anthropic's documentation read on 4 October 2026.
- Claude Code asks you to confirm a model switch only while the cache is still within its lifetime and the new model differs from the one that produced the last response. Before v2.1.238 it asked even after the cache had expired.
- Anthropic's sample for the PreModelSwitch hook shows 182,340 context tokens and an estimated cache write cost of $1.1396 for a switch from Sonnet 5 to Opus 5.
- Changing effort keeps the cache on Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription, and loses it on most other models.
- With the opusplan model setting, each change into or out of plan mode is a model switch that starts a new cache.
Claude Code is Anthropic's coding tool: you type a request in a terminal, and an AI model reads files, runs commands and edits code. Each message becomes a new request that contains the whole conversation again. The work is counted in tokens, the pieces of text that the model processes.
The prompt cache is a store of request text that the service has already processed. When the start of a new request matches the stored text exactly, the service reads that part from the store at a lower price. A request that finds its start in the store is a cache hit.
Two settings can make the whole stored copy unusable in one step: the model, and the effort level, which is the setting for how much the model reasons before it replies. This page explains what each change costs, when Claude Code warns you, and which changes keep the cache. It belongs to the guide to the Claude Code prompt cache. Which model to choose for which task is a separate question, answered on the related guide's page about which model and effort level to use.
Each model has its own cache
Anthropic's documentation, read on 4 October 2026, states the rule in two sentences: "Each model has its own cache. Switching with /model means the next request reads the entire conversation history with no cache hits, even though the content is identical." [1] /model is the command that changes the model during a session.
The text of your conversation is the same before and after the switch. The stored copy belongs to the old model, and the new model cannot read it. After that one request, the new model has its own stored copy [1].
When Claude Code asks you to confirm a switch
Claude Code warns you before a costly switch, under two conditions. The documentation says that when you run /model at the terminal, "Claude Code asks you to confirm the switch only while the cache is still warm and the new model isn't the one that produced the last response" [1].
Warm is Anthropic's word for a cache that is still within its lifetime. The lifetime has a technical name, TTL (time to live): how long a stored copy is kept when nobody uses it. The page says the cache stays warm for one TTL after Claude Code last sent a request in the conversation or Claude last responded. After that time the cache has expired, and Claude Code switches without asking [1]. The opposite of warm is cold: a cache whose stored copy has expired.
One version note applies. Before v2.1.238, Claude Code did not check the lifetime and asked even after the cache had expired [1].
The lifetime is five minutes or one hour. The default depends on how you are billed. The page on how long the cache lasts covers which one you get.
A hook that shows the cost before a switch
A hook is a script that Claude Code runs by itself at a fixed point. Anthropic's hooks reference, read on 4 October 2026, describes one hook for this exact moment. PreModelSwitch runs before Claude Code applies a model switch that you or a connected program requested, and it "can block it or ask you to confirm" [2] [3]. It requires Claude Code v2.1.251 or later [3].
The hook receives the facts you need to judge the cost [3]:
| Field the hook receives | What it tells you |
|---|---|
from_model and to_model |
The model the session is leaving and the model it is changing to |
context_tokens |
The tokens that the next request sends again as its prompt |
prompt_cache_warm |
Whether the current model's cache is likely still within its lifetime, which means the switch loses it |
cache_ttl |
The cache lifetime Claude Code requests for this session: 5m or 1h |
estimated_cache_write_usd |
The estimated cost in US dollars of writing context_tokens to the cache on the new model, without the next response |
pricing |
How that estimate was priced: configured (your organisation's own rates), catalog (list price) or default (an assumed rate for a model with no known price) |
The hook can answer in three ways. allow lets the switch proceed and skips Claude Code's own confirmation. deny cancels the switch. ask shows the user a confirmation with the hook's own message [3]. Only /model in an interactive session can show the ask prompt, and on every other surface Claude Code treats ask as a refusal [3]. A hook that does not respond within its time limit, 30 seconds by default, blocks the switch [3].
The hook runs for switches that are requested: /model, the model picker, the Model setting in /config, turning on fast mode when that changes the session's model, and a model change requested by a connected program such as an Agent SDK host or Remote Control. It does not run for switches that Claude Code makes by itself, such as an automatic model fallback [3]. Those are covered below.
A worked example with list prices
Anthropic's hooks reference includes a sample of what the hook receives when a user types /model opus in a session running Sonnet 5 [3]:
from_model:claude-sonnet-5to_model:claude-opus-5context_tokens: 182340prompt_cache_warm: truecache_ttl:5mestimated_cache_write_usd: 1.1396pricing:catalog
The estimate can be checked against Anthropic's pricing page, read on 4 October 2026 [4]. The page lists a five-minute cache write on Claude Opus 5 at $6.25 per million tokens. Our arithmetic: 182,340 tokens at $6.25 per million is $1.139625, which matches the sample's 1.1396.
Now compare the case where the user stays on Sonnet 5. The pricing page lists a cache read on Claude Sonnet 5 at $0.20 per million tokens [4]. Our arithmetic: if all 182,340 tokens were read from the cache on Sonnet 5, they would cost $0.036, to three decimal places. The switch replaces that with an estimated $1.14. Both figures leave out the price of the next response.
Anthropic attaches a caution to the field: the server may not need to store the whole context again, "so treat it as an estimate" [3]. The dollar figures are API list prices. Anthropic's pages give no statement of how cache tokens count against the usage limits of a subscription plan.
The second example is invented, to show the same sum on two current models. Suppose a conversation of 100,000 tokens. The prices are the list prices per million tokens on Anthropic's pricing page [4], and the costs are our arithmetic. The table assumes, as Anthropic's estimated_cache_write_usd field does, that after a switch the whole conversation is billed as a cache write on the new model.
| Case for a 100,000-token conversation (invented) | List price per million tokens | Cost of the stored text on the next request |
|---|---|---|
| Stay on Claude Sonnet 5.5: cache read | $0.20 | $0.02 |
| Stay on Claude Opus 5.5: cache read | $0.20 | $0.02 |
| Switch to Claude Sonnet 5.5: five-minute cache write | $2.50 | $0.25 |
| Switch to Claude Sonnet 5.5: one-hour cache write | $4 | $0.40 |
| Switch to Claude Opus 5.5: five-minute cache write | $5 | $0.50 |
| Switch to Claude Opus 5.5: one-hour cache write | $8 | $0.80 |
The cost of a switch grows with the size of the conversation. The same switch in a 10,000-token conversation would cost one tenth of each figure in the table.
Switches that happen without the /model command
Three other events are model switches in Anthropic's documentation. Each one has the same effect on the cache as typing /model.
The opusplan setting. Plan mode is a setting in which Claude researches and proposes changes without making them. opusplan is a model setting that uses Opus during plan mode and Sonnet during execution [2]. Anthropic's caching page says that each change into or out of plan mode is then a model switch and starts a new cache [1]. A session that enters plan mode once and leaves it once has two switches.
Automatic model fallback. Fable models, Opus 5.5, Sonnet 5.5 and Opus 5 run with safety classifiers, which are checks that flag certain content. The model configuration page says that when a classifier flags a request in a category that has a fallback model, Claude Code runs the request again on that model, and the session continues there [1] [2]. For example, a request flagged for cybersecurity content on Opus 5.5 runs again on Opus 4.8 [2]. Category-based fallback requires Claude Code v2.1.219 or later [2]. To decide each time yourself, set switchModelsOnFlag to false in your settings file. A flagged request then pauses the session and offers two options: switch to the fallback model, or edit the prompt and retry on the current model [2].
A skill or command that names a model. A skill is a packaged set of instructions for one task, stored in a file. The settings block at the start of that file is called frontmatter. When the frontmatter names a model other than the session's current model, "that turn is also a model switch: the next request reads the entire conversation history with no cache hits" [1]. The session model resumes on your next prompt [1]. One case is different. A subagent is a second copy of Claude that works on a task in its own separate conversation. In a skill marked context: fork, which runs in a subagent, the model value sets the forked subagent's model instead [1].
Effort level: which models keep the cache
The rule for effort has two parts, both from Anthropic's caching page, read on 4 October 2026 [1].
On most models, changing the effort level in the middle of a session "means the next request reads the entire conversation history with no cache hits". While the cache is still within its lifetime, Claude Code asks you to confirm the change first.
On Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription, changing effort keeps the cache, and Claude Code applies the new level without asking.
The second part has exceptions. It does not apply on Amazon Bedrock, on Google Cloud's Agent Platform or on a Claude apps gateway. It also does not apply when you set the environment variable CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS, or when your organisation has a HIPAA configuration [1]. In those cases an effort change on the three models has the same cost as on other models.
One version note: before v2.1.260, changing effort on Fable 5.1 with an API key or a Claude subscription also invalidated the cache, which means it made the stored copy unusable [1].
The API is the service that receives each request. The API reference, read the same day, describes the general behaviour for programs that call it directly. Changing the effort value invalidates the cached messages. On models that support per-message effort, an effort change sent as a system message inside the conversation leaves the stored start intact [6].
There is also a way to ask for more reasoning on one turn and keep the session's effort setting. The model configuration page says that if you include the word ultrathink anywhere in your prompt, Claude Code adds an instruction inside the conversation for that turn, and "the effort level sent to the API is unchanged" [2].
Fast mode: a header that is part of the cache identity
Fast mode is a setting that makes Claude Opus respond faster at a higher price per token. It is supported on Opus 5.5, Opus 5 and Opus 4.8 [5].
Turning it on adds a header to the request. A header is a labelled value sent with the request, outside the conversation text. Anthropic says this header is part of the cache key, the set of values the service uses to identify a stored copy. The first request sent with fast mode on therefore reads the entire conversation history with no cache hits [1].
Three details decide the cost [1] [5]:
- Claude Code sets the header once when a turn starts. If you turn fast mode on while Claude is working, the cache miss, a request that finds no stored copy, happens on the first request of your next turn.
- The fast mode page says that the first time you enable fast mode in a conversation, "you pay the full fast mode uncached input token price for the entire conversation context". The pricing page lists fast mode input on Claude Opus 5.5 at $8 per million tokens, against $4 at standard speed [4]. In an invented 100,000-token conversation, that is $0.80 by our arithmetic.
- The cost applies once per conversation. Turning fast mode off and on again later keeps the cache.
If your current model does not support fast mode, turning it on also switches your model to Opus, and that is a model switch with its own new cache [1] [5].
The changes in one table
| Change in the middle of a session | Cache | Where it applies |
|---|---|---|
/model to a different model |
Lost: the next request reads the entire history with no cache hits | All models |
Entering or leaving plan mode with opusplan |
Lost on each change | Sessions with the opusplan model setting |
| Automatic model fallback | Lost: the session continues on the fallback model | Fable models, Opus 5.5, Sonnet 5.5 and Opus 5 |
| A skill or command whose frontmatter names another model | Lost for that turn | Any session. With context: fork the value sets the subagent's model instead |
| Changing effort level | Kept | Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription (Fable 5.1 from v2.1.260) |
| Changing effort level | Lost | Most other models. Also the three models above on Amazon Bedrock, Google Cloud's Agent Platform or a Claude apps gateway, with CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS set, or with a HIPAA configuration |
| Turning on fast mode for the first time | Lost once per conversation | Opus 5.5, Opus 5 and Opus 4.8 |
| Turning fast mode off, or on again later | Kept | The same three models |
Our position
Choose the model and the effort level before the first message. That is Anthropic's own tip on the caching page: choose both at the start of a session, because "the fewer changes you make mid-task, the higher your cache hit rate" [1]. The cache hit rate is the share of your input that is read from the store.
When a change in the middle of a task is worth making, Reveneau recommends four checks, each based on the facts above.
- Look at the size of the conversation. Anthropic's estimate field prices a switch as the whole conversation written to the new model's cache. The switch costs less in a short conversation, for example after
/clearhas started a new one. - Read the confirmation. Claude Code asks only while the cache is still within its lifetime and the new model differs from the one that produced the last response. When it asks, the switch will lose a usable cache.
- Use
ultrathinkfor one hard step instead of raising the effort level on a model where that change loses the cache. - Give a side task to a subagent when it needs a different model. Anthropic's page says a subagent builds a cache of its own and the parent's cache is unaffected [1]. The page on actions that keep the cache covers subagents.
For a team, a PreModelSwitch hook that returns ask shows the developer the token count before each switch. The full list of actions with a cache cost is on the page about the nine actions that reset the cache.
Reveneau is an AI software development consultancy. All of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. The prices and product statements on this page are Anthropic's own, read on 4 October 2026, and each sum marked as ours is arithmetic on those prices.
Common questions
Does switching models in Claude Code reset the prompt cache?
Yes. Switching models in Claude Code means the next request reads the entire conversation history with no cache hits. Anthropic's documentation, read on 4 October 2026, gives the reason: each model has its own cache, so the stored copy made on the old model cannot be read by the new one, even though the content is identical. After that one request, the new model has its own stored copy.
Why does Claude Code ask me to confirm a model switch?
Claude Code asks you to confirm a model switch as a warning before a costly change. Anthropic's documentation, read on 4 October 2026, says the confirmation appears only while the cache is still warm, meaning within its lifetime, and when the new model differs from the one that produced the last response. The cache stays warm for one lifetime period after the last request Claude Code sent in the conversation or the last response from Claude.
Why did Claude Code switch models without asking me?
Claude Code switches models without asking when the cache has already expired. Anthropic's documentation, read on 4 October 2026, says the cache stays warm, meaning within its lifetime, for one lifetime period after the last request or response, and once that time passes Claude Code switches without asking. Before v2.1.238, Claude Code did not check the lifetime and asked even after the cache had expired. A `PreModelSwitch` hook that returns `allow` also skips the confirmation.
What is a PreModelSwitch hook?
A `PreModelSwitch` hook is a script that Claude Code runs before it applies a model switch that you or a connected program requested. Anthropic's hooks reference, read on 4 October 2026, says it requires Claude Code v2.1.251 or later and can allow the switch, deny it, or ask the user to confirm. The hook receives the token count of the context and an estimated cache write cost in US dollars.
How much does a model switch cost in dollars?
The cost of a model switch in dollars depends on the size of the conversation and the new model's cache write price. In the sample in Anthropic's hooks reference, read on 4 October 2026, a switch from Sonnet 5 to Opus 5 with 182,340 context tokens has an estimated cache write cost of $1.1396. By our arithmetic, that matches Anthropic's list price of $6.25 per million tokens for a five-minute cache write on Claude Opus 5.
Does opusplan lose the cache when I leave plan mode?
Yes. With the `opusplan` model setting, leaving plan mode is a model switch, and so is entering it. Anthropic's documentation, read on 4 October 2026, says `opusplan` uses Opus during plan mode and Sonnet during execution, and that each plan-mode change is a model switch that starts a new cache. A session that enters plan mode once and leaves it once therefore has two model switches.
Does automatic model fallback reset the prompt cache?
Yes. Anthropic's documentation, read on 4 October 2026, counts automatic model fallback as a model switch, and each model has its own cache. Fallback happens on Fable models, Opus 5.5, Sonnet 5.5 and Opus 5 when a safety classifier flags a request in a category that has a fallback model. Claude Code runs the request again on that model and the session continues there. With `switchModelsOnFlag` set to `false`, a flagged request pauses the session and offers you the choice.
Can a skill cause a model switch in Claude Code?
Yes. A skill or command causes a model switch when its frontmatter, the settings block at the start of its file, names a model other than the session's current model. Anthropic's documentation, read on 4 October 2026, says the next request then reads the entire conversation history with no cache hits, and the session model resumes on your next prompt. In a skill marked `context: fork`, the value sets the forked subagent's model instead.
Does changing the effort level reset the prompt cache on Opus 5.5?
No, when you use an API key or a Claude subscription. Anthropic's documentation, read on 4 October 2026, says that on Opus 5.5, Sonnet 5.5 and Fable 5.1, changing the effort level keeps the cache, and Claude Code applies the new level without asking. On most other models, an effort change in the middle of a session means the next request reads the entire conversation history with no cache hits.
Does changing the effort level reset the cache on Amazon Bedrock?
Yes. On Amazon Bedrock, an effort change on Opus 5.5, Sonnet 5.5 or Fable 5.1 has the same cost as on other models. Anthropic's documentation, read on 4 October 2026, says the rule that keeps the cache on those three models does not apply on Amazon Bedrock, on Google Cloud's Agent Platform or on a Claude apps gateway. It also does not apply with `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS` set or with a HIPAA configuration.
Does the ultrathink keyword change the effort level?
No. The `ultrathink` keyword asks for more reasoning on one turn and keeps the session's effort setting as it is. Anthropic's model configuration page, read on 4 October 2026, says that when the word appears anywhere in your prompt, Claude Code adds an instruction inside the conversation, and the effort level sent to the API is unchanged. Reveneau recommends it for one hard step on a model where an effort change loses the cache.
When does a model switch cost the least?
A model switch costs the least at the start of a session, when the conversation is shortest. Anthropic's estimate field prices a switch as the whole conversation written to the new model's cache, so the cost grows with the size of the conversation. In our invented example, a switch in a 10,000-token conversation costs one tenth of the same switch in a 100,000-token conversation. Anthropic's tip, read on 4 October 2026, is to choose the model at the start of a session.
References
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Model configuration (code.claude.com), read 4 October 2026
- Anthropic, Hooks reference (code.claude.com), read 4 October 2026
- Anthropic, Pricing (platform.claude.com), read 4 October 2026
- Anthropic, Speed up responses with fast mode (code.claude.com), read 4 October 2026
- Anthropic, Prompt caching (platform.claude.com), read 4 October 2026
More in What breaks it
The nine actions that reset the Claude Code prompt cache
Nine actions reset the Claude Code prompt cache, by Anthropic's documentation read on 4 October 2026: switching models, changing effort level, turning on fast mode, connecting or removing an MCP server, enabling or disabling a plugin, denying an entire tool, compacting the conversation, accumulating many images, and upgrading Claude Code. Each one can make the next request miss part or all of the stored text, which Anthropic describes as a one-time slower, more expensive turn. A model switch recomputes the whole request. An MCP server change and a tool deny rule keep the cache while tool search is on. This page lists what each action recomputes, the version notes, and how to avoid the cost.
What /compact costs while the cache is active and after it has expired
/compact costs a fraction of what the conversation size suggests while the prompt cache is still active, and it costs the most after the cache has expired. Anthropic's documentation, read on 4 October 2026, explains why: to write the summary, Claude Code sends a separate request with the same system prompt, tools and history as your conversation, plus one instruction at the end. While the cache is still within its lifetime, that request reads the stored start from the cache. After a break longer than the cache lifetime, it reprocesses the full history as uncached input. In both cases the turn after compaction stores only the shorter history. This page compares the two cases and the alternatives.