Why switching model in the middle of a Claude Code task makes the next turn slower and more expensive

You are an hour into a task in Claude Code on Sonnet. One step needs more reasoning, so you type /model opus. The next turn takes longer than any turn before it. Nothing went wrong. This post explains why, from the one rule in Anthropic's documentation that decides it, and then gives the moments when a switch costs the least.
Three terms first. Claude Code is Anthropic's coding tool: you type a request in a terminal, the text window where you type commands, and an AI model reads files, runs commands and edits code for you. The work is counted in tokens, which are pieces of text the model processes. The prompt cache is a store of request text that the service running the model has already processed. When the start of a new request matches stored text, that part is billed at a lower price.
The cache matches the start of a request, exactly
Claude Code sends the whole conversation with every request, by Anthropic's documentation read on 4 October 2026: the system prompt, your project context, every earlier message and tool result, and your new message. New content goes at the end, so most of each request is identical to the one before it. The service compares the start of a new request, which Anthropic calls the prefix, with text it processed recently, and reuses its earlier work where they match.
One sentence in Anthropic's documentation on prompt caching decides everything that follows: "The match is exact, so a change anywhere in the prefix recomputes everything after it." The same page adds that there is no caching per file or per segment. The stored copy covers the start of the request as one piece. Change one thing early, and everything after it is processed again.
The prices show why this matters. Anthropic's pricing page lists a cache read at 0.1 times the base input price on most models, and a cache write, which is text stored for the first time, at 1.25 times when the copy is kept for five minutes. On Claude Sonnet 5.5, that is $0.20 per million tokens for a read against $2.50 for a write.
Each model has its own cache
Anthropic's rule for a switch is two sentences: "Each model has its own cache. Switching with /model means the next request reads the entire conversation history with no cache hits, even though the content is identical."
The stored copy belongs to the old model, and the new model cannot read it. So the first request on the new model is processed in full, and billed as a cache write at the new model's price. After that one request, the new model has its own stored copy. Anthropic describes the result as "a one-time slower, more expensive turn, after which the new prefix is cached".
Anthropic's hooks reference gives a sample of the size. A hook is a script that Claude Code runs by itself at a fixed point, and the PreModelSwitch hook, which requires Claude Code v2.1.251 or later, runs before a requested model switch and receives the facts needed to judge it. In Anthropic's sample, a user types /model opus in a session running Sonnet 5. The hook receives 182,340 context tokens and an estimated cache write cost of $1.1396, priced at list price. Our arithmetic checks it: Anthropic's pricing page lists a five-minute cache write on Claude Opus 5 at $6.25 per million tokens, and 182,340 tokens at that price is $1.139625. Reading the same tokens from the cache on Sonnet 5, at the listed $0.20 per million, would cost $0.036. The switch replaces four cents with $1.14. Anthropic adds that the server may not need to store the whole context again, so the field is an estimate.
An invented example shows the same sum on two current models. In a conversation of 100,000 tokens, reading it from the cache on Claude Sonnet 5.5 or Claude Opus 5.5 costs $0.02. Switching to Sonnet 5.5 writes it to the five-minute cache at $2.50 per million tokens, which is $0.25, and switching to Opus 5.5 writes it at $5 per million, which is $0.50. The page on switching model or effort level in the middle of a task has the full table.
When Claude Code asks you to confirm a switch
Claude Code asks you to confirm a switch under two conditions, by Anthropic's documentation: only while the cache is still within its lifetime, and only when the new model differs from the one that produced the last response. The lifetime is one hour on a subscription within plan usage and five minutes with usage credits, an API key or a cloud provider. After it has passed, there is no stored copy left to lose, and Claude Code switches without asking. Before v2.1.238, it asked even after the cache had expired. When the prompt appears, the switch will lose a usable cache.
Three other events are model switches in Anthropic's documentation, with the same effect as typing /model: the opusplan setting, which uses Opus in plan mode and Sonnet afterwards, so each change into or out of plan mode starts a new cache; automatic model fallback, which runs a flagged request again on another model; and a skill, a packaged set of instructions for one task, that names a different model in its settings. The page on the nine actions that reset the cache lists every one, with version notes.
The effort level, which is the setting for how much the model reasons before it replies, follows a split rule. On most models, changing it in the middle of a session means the next request reads the whole history with no cache hits. On Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription, changing effort keeps the cache.
What we recommend
Anthropic's own tip is one sentence: "Pick your model and effort level at the top of a session, then save /compact for natural breaks between tasks. The fewer changes you make mid-task, the higher your cache hit rate." The cache hit rate is the share of your input that is read from the store. We add three habits that follow from the facts above.
- Switch after
/clear, when the conversation is empty. The cost of a switch is set by the length of the conversation at that moment. - Use
ultrathinkfor one hard step. Anthropic's model configuration page says that when the word appears anywhere in your prompt, Claude Code adds an instruction for that turn and the effort level sent to the API is unchanged. The cache is kept. - Give a side task that needs a different model to a subagent. A subagent is a second copy of Claude that works in its own context window, with a cache of its own, and Anthropic's page says the parent's cache is unaffected.
To see whether a switch cost you, run /usage. With Claude Code v2.1.251 or later, its Prompt cache (main) line shows the share of input tokens read from the cache and the number of misses. The page on how to check your cache hit rate reads that line field by field.
Reveneau is an AI software development consultancy. All of its code is written by AI, so token use is a running cost of every Reveneau build, and the guide to the Claude Code prompt cache collects every action that keeps or resets it. Reveneau is independent of Anthropic. Every price and product statement in this post is Anthropic's own, read on 4 October 2026, and each sum marked as ours is arithmetic on those prices.
A new model reads the conversation from the first page, and that first reading is what a switch costs.
Sources
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026: the exact-match rule, each model having its own cache, the confirmation conditions, the effort rule,
opusplan, fast mode, and the tip. - Anthropic, Hooks reference (code.claude.com), read 4 October 2026: the
PreModelSwitchhook and its sample of 182,340 tokens and $1.1396. - Anthropic, Pricing (platform.claude.com), read 4 October 2026: the cache read and write multipliers and every list price in this post.
- Anthropic, Model configuration (code.claude.com), read 4 October 2026:
ultrathinkand automatic model fallback. - Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026: the
Prompt cache (main)line in/usage.


