Engineering

Why switching model in the middle of a Claude Code task makes the next turn slower and more expensive

Editorial · Reveneau · October 6, 2026

Why switching model in the middle of a Claude Code task makes the next turn slower and more expensive

You are an hour into a task in Claude Code on Sonnet. One step needs more reasoning, so you type /model opus. The next turn takes longer than any turn before it. Nothing went wrong. This post explains why, from the one rule in Anthropic's documentation that decides it, and then gives the moments when a switch costs the least.

Three terms first. Claude Code is Anthropic's coding tool: you type a request in a terminal, the text window where you type commands, and an AI model reads files, runs commands and edits code for you. The work is counted in tokens, which are pieces of text the model processes. The prompt cache is a store of request text that the service running the model has already processed. When the start of a new request matches stored text, that part is billed at a lower price.

The cache matches the start of a request, exactly

Claude Code sends the whole conversation with every request, by Anthropic's documentation read on 4 October 2026: the system prompt, your project context, every earlier message and tool result, and your new message. New content goes at the end, so most of each request is identical to the one before it. The service compares the start of a new request, which Anthropic calls the prefix, with text it processed recently, and reuses its earlier work where they match.

One sentence in Anthropic's documentation on prompt caching decides everything that follows: "The match is exact, so a change anywhere in the prefix recomputes everything after it." The same page adds that there is no caching per file or per segment. The stored copy covers the start of the request as one piece. Change one thing early, and everything after it is processed again.

The prices show why this matters. Anthropic's pricing page lists a cache read at 0.1 times the base input price on most models, and a cache write, which is text stored for the first time, at 1.25 times when the copy is kept for five minutes. On Claude Sonnet 5.5, that is $0.20 per million tokens for a read against $2.50 for a write.

Each model has its own cache

Anthropic's rule for a switch is two sentences: "Each model has its own cache. Switching with /model means the next request reads the entire conversation history with no cache hits, even though the content is identical."

The stored copy belongs to the old model, and the new model cannot read it. So the first request on the new model is processed in full, and billed as a cache write at the new model's price. After that one request, the new model has its own stored copy. Anthropic describes the result as "a one-time slower, more expensive turn, after which the new prefix is cached".

Anthropic's hooks reference gives a sample of the size. A hook is a script that Claude Code runs by itself at a fixed point, and the PreModelSwitch hook, which requires Claude Code v2.1.251 or later, runs before a requested model switch and receives the facts needed to judge it. In Anthropic's sample, a user types /model opus in a session running Sonnet 5. The hook receives 182,340 context tokens and an estimated cache write cost of $1.1396, priced at list price. Our arithmetic checks it: Anthropic's pricing page lists a five-minute cache write on Claude Opus 5 at $6.25 per million tokens, and 182,340 tokens at that price is $1.139625. Reading the same tokens from the cache on Sonnet 5, at the listed $0.20 per million, would cost $0.036. The switch replaces four cents with $1.14. Anthropic adds that the server may not need to store the whole context again, so the field is an estimate.

An invented example shows the same sum on two current models. In a conversation of 100,000 tokens, reading it from the cache on Claude Sonnet 5.5 or Claude Opus 5.5 costs $0.02. Switching to Sonnet 5.5 writes it to the five-minute cache at $2.50 per million tokens, which is $0.25, and switching to Opus 5.5 writes it at $5 per million, which is $0.50. The page on switching model or effort level in the middle of a task has the full table.

When Claude Code asks you to confirm a switch

Claude Code asks you to confirm a switch under two conditions, by Anthropic's documentation: only while the cache is still within its lifetime, and only when the new model differs from the one that produced the last response. The lifetime is one hour on a subscription within plan usage and five minutes with usage credits, an API key or a cloud provider. After it has passed, there is no stored copy left to lose, and Claude Code switches without asking. Before v2.1.238, it asked even after the cache had expired. When the prompt appears, the switch will lose a usable cache.

Three other events are model switches in Anthropic's documentation, with the same effect as typing /model: the opusplan setting, which uses Opus in plan mode and Sonnet afterwards, so each change into or out of plan mode starts a new cache; automatic model fallback, which runs a flagged request again on another model; and a skill, a packaged set of instructions for one task, that names a different model in its settings. The page on the nine actions that reset the cache lists every one, with version notes.

The effort level, which is the setting for how much the model reasons before it replies, follows a split rule. On most models, changing it in the middle of a session means the next request reads the whole history with no cache hits. On Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription, changing effort keeps the cache.

What we recommend

Anthropic's own tip is one sentence: "Pick your model and effort level at the top of a session, then save /compact for natural breaks between tasks. The fewer changes you make mid-task, the higher your cache hit rate." The cache hit rate is the share of your input that is read from the store. We add three habits that follow from the facts above.

  1. Switch after /clear, when the conversation is empty. The cost of a switch is set by the length of the conversation at that moment.
  2. Use ultrathink for one hard step. Anthropic's model configuration page says that when the word appears anywhere in your prompt, Claude Code adds an instruction for that turn and the effort level sent to the API is unchanged. The cache is kept.
  3. Give a side task that needs a different model to a subagent. A subagent is a second copy of Claude that works in its own context window, with a cache of its own, and Anthropic's page says the parent's cache is unaffected.

To see whether a switch cost you, run /usage. With Claude Code v2.1.251 or later, its Prompt cache (main) line shows the share of input tokens read from the cache and the number of misses. The page on how to check your cache hit rate reads that line field by field.

Reveneau is an AI software development consultancy. All of its code is written by AI, so token use is a running cost of every Reveneau build, and the guide to the Claude Code prompt cache collects every action that keeps or resets it. Reveneau is independent of Anthropic. Every price and product statement in this post is Anthropic's own, read on 4 October 2026, and each sum marked as ours is arithmetic on those prices.

A new model reads the conversation from the first page, and that first reading is what a switch costs.

Sources

Common questions

Why is the first turn after a model switch slower in Claude Code?

The first turn after a model switch is slower because the new model has no stored copy of the conversation. Anthropic's documentation, read on 4 October 2026, says each model has its own prompt cache, so after a switch with /model the next request reads the entire conversation history with no cache hits, even though the content is identical. The service processes every token again, which takes longer and is billed at the new model's cache write price.

What is the exact-match rule for the Claude Code prompt cache?

The exact-match rule means the prompt cache compares the start of a new request with stored text and reuses it only when the two are identical. Anthropic's documentation, read on 4 October 2026, says the match is exact, so a change anywhere in that start recomputes everything after it, and that there is no caching per file or per segment. A different model is a different cache altogether, so a switch leaves nothing to match.

How much does a model switch cost in Claude Code?

The cost depends on the size of the conversation and the new model's cache write price. Anthropic's hooks reference, read on 4 October 2026, gives a sample: a switch from Sonnet 5 to Opus 5 with 182,340 context tokens has an estimated cache write cost of $1.1396. By our arithmetic that equals Anthropic's listed five-minute cache write price on Claude Opus 5, $6.25 per million tokens. Anthropic says to treat the field as an estimate.

Why does Claude Code ask me to confirm a model switch sometimes and not other times?

Claude Code asks you to confirm a model switch only while the cache is still within its lifetime and the new model differs from the one that produced the last response, by Anthropic's documentation read on 4 October 2026. Once the lifetime has passed, there is no stored copy to lose, so Claude Code switches without asking. Before v2.1.238 it asked even after the cache had expired. A confirmation means the switch will lose a usable cache.

Does changing the effort level reset the Claude Code cache?

On Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription, changing the effort level keeps the cache, by Anthropic's documentation read on 4 October 2026, and Fable 5.1 joined that list in v2.1.260. On most other models, and on Amazon Bedrock, Google Cloud's Agent Platform or a Claude apps gateway, an effort change has the same cost as a model switch. The effort level is the setting for how much the model reasons before it replies.

Does opusplan lose the cache every time I enter plan mode?

Yes. The opusplan model setting uses Opus during plan mode and Sonnet while Claude writes code, and Anthropic's documentation, read on 4 October 2026, says each change into or out of plan mode is a model switch that starts a new cache. Plan mode is the setting in which Claude proposes an approach and edits nothing until you approve. A session that enters plan mode once and leaves it once has two switches, so opusplan suits a task you plan once at the start.

Does turning on fast mode reset the prompt cache?

Yes, once per conversation. Fast mode adds a header to each request, a labelled value sent outside the conversation text, and Anthropic's documentation, read on 4 October 2026, says the header is part of the cache key. The first request with fast mode on therefore reads the whole conversation with no cache hits, at the fast mode input price, which Anthropic's pricing page lists at $8 per million tokens on Claude Opus 5.5. Turning it off and on again later keeps the cache.

When is the cheapest moment to switch models in Claude Code?

The cheapest moment to switch models is when the conversation is shortest, which means before the first message or right after /clear. Anthropic's estimate field prices a switch as the whole conversation written to the new model's cache, so the cost grows with the conversation. In an invented example on Anthropic's list prices read on 4 October 2026, a switch to Opus 5.5 in a 100,000-token conversation costs $0.50 as a five-minute cache write, and one tenth of that at 10,000 tokens.

How can I get more reasoning for one step without switching model or effort?

Include the word ultrathink in the request. Anthropic's model configuration page, read on 4 October 2026, says that when the word appears anywhere in your prompt, Claude Code adds an instruction inside the conversation for that turn and the effort level sent to the API is unchanged, so the cache is kept. For a side task that needs a different model, give it to a subagent, a second copy of Claude with its own context and its own cache, which leaves the main conversation's cache in place.