How it works

What the Claude Code prompt cache is and why every message re-sends everything

The Claude Code prompt cache is a store, kept by the service that runs the model, of request text it has already processed. Claude Code needs it because the model keeps nothing between requests, so every message re-sends the whole conversation. When the start of a new request matches the stored text exactly, that part is billed as a cache read. On Anthropic's pricing page, read on 4 October 2026, a cache read costs 0.1 times the base input price on most models, and a five-minute cache write costs 1.25 times. Claude Code manages the cache automatically. This page explains the match rule, the token types and where the cache is stored.

Published October 4, 2026. Editorial.

Key takeaways

  • Claude Code re-sends the system prompt, the project context and every earlier message with each request, because the model keeps nothing between requests, by Anthropic's documentation read on 4 October 2026.
  • The prompt cache matches the start of a request exactly, and Anthropic's documentation says a change anywhere in that start recomputes everything after it.
  • On Anthropic's pricing page, read on 4 October 2026, a cache read costs 0.1 times the base input price on most models, 0.05 times on Claude Opus 5.5 and 0.025 times on Claude Fable 5.1.
  • A five-minute cache write costs 1.25 times the base input price and a one-hour cache write costs 2 times, on the same pricing page.
  • In Anthropic's sample of the /usage screen, a Claude Sonnet 4.6 session shows 940.0k cache read tokens beside 1.2k input tokens and 50.0k cache write tokens.

Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. The work is counted in tokens. A token is a piece of text that the model processes. Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters or 0.75 of an English word [3].

The prompt cache is a store of request text that the service running the model has already processed. When a new request starts with the same text, the service reuses its earlier work on that text and charges a lower price for it. This page explains why Claude Code needs that store, how the match works, and what each type of token costs. It is the first page of the guide to the Claude Code prompt cache.

Why every message sends the whole conversation again

The model keeps nothing between two requests. Anthropic's documentation, read on 4 October 2026, says: "The model doesn't remember anything between requests, so Claude Code re-sends the full context: the system prompt, your project context, every prior message and tool result, and your new message." [1]

Each of those items needs one plain explanation:

  • The system prompt is the set of instructions that Anthropic writes for the model. In Anthropic's description it holds the core instructions and the tool definitions, which describe each action the model can ask Claude Code to perform [1].
  • Your project context is the text Claude Code loads from your project at the start of a session. It includes CLAUDE.md, a file of instructions that you write for Claude, and auto memory, the notes Claude saved for itself in earlier sessions [1].
  • A tool result is what comes back when Claude reads a file or runs a command.

A turn is one message from you plus the work Claude does in reply. One turn can be several requests. Anthropic's page on costs says that each time Claude uses tools, Claude Code sends another request that contains that group of tool results [4].

New content is added at the end of the request. The documentation states the result: "most of each request is identical to the one before it" [1]. The API is the service that receives each request. Without caching, the same page says, the API would reprocess your full history on every turn [1].

The related guide's page on where Claude Code tokens go lists what is in a request and how large each part is.

How the match works: the unchanged start of the request

The cache works on the start of the request. Anthropic's name for that start is the prefix. In the words of the documentation: "The API caches by matching the start of each request, called the prefix, against content it recently processed." [1] On a normal turn, the prefix is the entire previous request, and only the latest exchange is new [1].

Four rules follow. All four come from Anthropic's pages, read on 4 October 2026.

The match is exact. The Claude Code page says: "The match is exact, so a change anywhere in the prefix recomputes everything after it." [1] The API reference states the same rule in its own terms. A cache hit is a request that finds a stored copy of its start. A hit requires "100% identical prompt segments, including all text and images" up to the marked point [2]. A cache miss is the opposite: a request that finds no stored copy for some or all of its start.

The store covers the whole start, in one fixed order. The Claude Code page says: "There is no per-file or per-segment caching." [1] The API reference gives the order in which a request is read: tool definitions first, then the system prompt, then the messages [2]. A file that Claude read is part of the messages. The service stores it as part of everything that came before it in the request.

A stored copy expires when nobody uses it. By default a stored copy lasts 5 minutes, and each use starts the 5 minutes again at no extra charge [2]. Anthropic also offers a one-hour lifetime at a higher write price [2]. The API reference adds that there is currently no way to clear the cache yourself: stored prefixes expire automatically after a minimum of 5 minutes without use [2]. The page on how long the cache lasts covers which lifetime Claude Code requests for you.

A start below a minimum size is processed without caching. The API returns no error in that case [2]. On the Claude API the minimum is 512 tokens for a group of models that includes Claude Fable 5.1, Claude Opus 5.5, Claude Opus 5 and Claude Sonnet 5.5. It is 1,024 tokens for a group that includes Claude Sonnet 5 and Claude Sonnet 4.6, and 4,096 tokens for Claude Haiku 4.5 [2].

What a cache read and a cache write are

The API reports token counts on every response. Anthropic's documentation names the two that show how the cache is working [1]:

  • A cache write is text stored for the first time on this turn. The API reports it as cache_creation_input_tokens, and it is billed at the cache write price.
  • A cache read is stored text used again on this turn. The API reports it as cache_read_input_tokens, and it is billed at the cached token price, which is below the standard input price.

A third count, input_tokens, holds the input that was neither read from the store nor written to it. The API reference says that the total input of a request is the three counts added together [2].

The table gives each price as a multiple of the base input price. The multipliers come from Anthropic's pricing page, read on 4 October 2026 [3].

Token type What it is Price relative to base input
Base input Text the model reads that is neither read from the store nor written to it 1 time
Cache write, five-minute lifetime Text stored for reuse for the first time 1.25 times
Cache write, one-hour lifetime The same text, kept for one hour 2 times
Cache read Stored text used again 0.1 times on most models, 0.05 times on Claude Opus 5.5, 0.025 times on Claude Fable 5.1 and Claude Mythos 5.1
Output Text the model writes Listed as a separate price for each model

The same page lists these API prices in US dollars per million tokens for two current models [3]. They are list prices for use that is billed per token. Anthropic's pages give no statement of how cache tokens count against the usage limits of a subscription plan.

Model Base input Five-minute cache write One-hour cache write Cache read Output
Claude Sonnet 5.5 $2 $2.50 $4 $0.20 $10
Claude Opus 5.5 $4 $5 $8 $0.20 $20

Two statements on Anthropic's pages explain how to read these numbers. The pricing page says that where a cache read costs 10 percent of the input price, one cache read is enough to recover the extra cost of a five-minute write, and it takes two cache reads to recover the extra cost of a one-hour write [3]. The API reference says that prompt caching has no effect on the output: "The response you receive is identical to what you would get if prompt caching were not used." [2]

Two examples with numbers

The first example is Anthropic's own. Its page on costs shows a sample of the /usage screen, the screen that reports the token counts of a session. The sample session ran on Claude Sonnet 4.6 and shows 1.2k input tokens, 5.3k output tokens, 940.0k cache read tokens and 50.0k cache write tokens, for a total of $0.55 [4].

The rest of this paragraph is our own arithmetic on that sample. The three input-side counts add up to 991,200 tokens, and the 940,000 cache read tokens are 94.8 percent of them. The pricing page lists Claude Sonnet 4.6 at $3 per million base input tokens and $0.30 per million cache read tokens [3]. At the cache read price, 940,000 tokens cost $0.282. At the base input price, the same 940,000 tokens would cost $2.82.

The second example is invented, and it uses the list prices for Claude Sonnet 5.5 from the table above. Suppose the stored start of a conversation is 100,000 tokens. The example assumes that a start which no longer matches is billed in full as a five-minute cache write. That is our assumption: Anthropic's page says such a request recomputes the full input and stores it again, and it gives no general price for that request [1].

  • On a turn where the start matches, the 100,000 tokens are cache reads: 100,000 tokens at $0.20 per million is $0.02.
  • On a turn where the start no longer matches, the example bills the 100,000 tokens as a new cache write: 100,000 tokens at the five-minute write price of $2.50 per million is $0.25.

Under that assumption, the second turn costs 12.5 times the first for the same text. With the Claude Opus 5.5 prices, the same invented conversation costs $0.02 to read and $0.50 to write again, which is 25 times. The conversation size is invented. The prices are Anthropic's list prices.

Claude Code manages the cache for you

You set up nothing. Anthropic's documentation, read on 4 October 2026, says: "Claude Code handles prompt caching for you, unless you disable it." [1] Two parts of that work are described on the same page.

First, Claude Code orders each request so that content which rarely changes between turns comes first: the system prompt, then the project context, then the conversation [1]. The page on the three layers of a request explains what is in each part.

Second, Claude Code sends markers with each request, named cache_control markers, which tell the API where the reusable start ends [1] [2].

Your own actions still matter. To invalidate the cache means to make the stored copy unusable for the next request. Anthropic's page says that some actions do this "and make the next response slower and more expensive while it rebuilds" [1]. The page on the nine actions that reset the cache lists them.

Anthropic gives one piece of advice in a tip on the same page. Choose your model and your effort level, the setting for how much the model reasons, at the start of a session. Then keep /compact, the command that replaces the conversation history with a summary, for natural breaks between tasks. The tip ends: "The fewer changes you make mid-task, the higher your cache hit rate." [1] The cache hit rate is the share of your input that is read from the store.

You can turn caching off. Setting the environment variable DISABLE_PROMPT_CACHING to 1 disables it for all models. Anthropic describes this as occasionally useful when debugging, and adds: "For normal use, leave caching enabled." [1]

Where the cache is stored

The cache is kept on the server side. Anthropic's page says that caching happens "in whichever infrastructure serves your model", and that the place depends on how you sign in [1]:

  • API key, Claude subscription or Claude Platform on AWS: the cache is in Anthropic's infrastructure.
  • Amazon Bedrock or Google Cloud's Agent Platform: the cache is in your cloud provider's serving infrastructure.
  • Microsoft Foundry: it depends on the hosting option of the deployment. Deployments hosted on Azure are served on Azure infrastructure, and deployments hosted on Anthropic are served on Anthropic's infrastructure.
  • A custom ANTHROPIC_BASE_URL or an LLM gateway: a gateway is a server that receives your requests first and passes them on. The cache is wherever your requests are forwarded, "and whether caching works depends on the gateway" [1]. The page on gateways and cloud providers covers this case.

The API reference describes who can read a stored copy. Caches are isolated between organisations: "Different organizations never share caches, even if they use identical prompts." [2] On the Claude API, Claude Platform on AWS and Microsoft Foundry, caches are also isolated between the workspaces inside one organisation [2]. The same reference says that Anthropic does not store the raw text of your prompts or Claude's responses for prompt caching, and that the cached representations are held in memory only [2].

Our position

Leave caching on, and plan a session so that the start of the request stays the same. The price table gives the reason: on most models a cache read costs one tenth of the base input price, and a five-minute cache write costs 1.25 times the base input price [3]. In a session that keeps its start unchanged, the stored text is billed at the first price on each turn. Reveneau recommends following Anthropic's tip: choose the model and the effort level before the first message, and run /compact between tasks.

Then check the result. The page on how to check your cache hit rate shows where Claude Code reports cache reads and cache writes.

Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. Every figure on this page is Anthropic's own statement about its own product, and the sums marked as ours are arithmetic on those figures.

Common questions

What is the prompt cache in Claude Code?

The prompt cache in Claude Code is a store of request text that the service running the model has already processed. When a new request starts with exactly the same text, the service reuses its earlier work and bills that text as a cache read. Anthropic's documentation, read on 4 October 2026, says Claude Code handles prompt caching for you unless you disable it, so you set up nothing.

What is a prefix in prompt caching?

A prefix in prompt caching is the start of a request, the part that the service compares with text it processed recently. Anthropic's documentation, read on 4 October 2026, says that on a normal turn the prefix is the entire previous request and only the latest exchange is new. The API reference gives the order in which a request is read: tool definitions, then the system prompt, then the messages.

How exact does the match have to be for the prompt cache to work?

The match has to be exact for the prompt cache to work. Anthropic's API reference, read on 4 October 2026, says a cache hit requires 100% identical prompt segments, including all text and images, up to the marked point. The Claude Code documentation states what follows from that: a change anywhere in the start of the request recomputes everything after it.

Is the cache read discount the same on every Claude model?

No. The cache read price differs between Claude models. On Anthropic's pricing page, read on 4 October 2026, a cache read costs 0.1 times the base input price on most models, 0.05 times on Claude Opus 5.5, and 0.025 times on Claude Fable 5.1 and Claude Mythos 5.1. In dollars, the page lists a cache read at $0.20 per million tokens on both Claude Sonnet 5.5 and Claude Opus 5.5.

How many cache reads does it take to recover the cost of a cache write?

It takes one cache read to recover the extra cost of a five-minute cache write, by Anthropic's pricing page, read on 4 October 2026. The page gives the reason: a five-minute write costs 1.25 times the base input price, and a cache read costs 10 percent of the input price on the standard multiplier. With the one-hour lifetime, where a write costs 2 times the base input price, the page says it takes two cache reads.

Do I need to turn on prompt caching in Claude Code?

No. Prompt caching is already on in Claude Code, and Claude Code manages it for you. Anthropic's documentation, read on 4 October 2026, says Claude Code orders each request so that content which rarely changes comes first, and it sends the `cache_control` markers itself. You can turn caching off by setting `DISABLE_PROMPT_CACHING` to `1`, which Anthropic describes as occasionally useful when debugging. For normal use, Anthropic says to leave caching enabled.

Where is the Claude Code prompt cache stored?

The Claude Code prompt cache is stored on the server side, in the infrastructure that serves your model. By Anthropic's documentation, read on 4 October 2026, that is Anthropic's infrastructure with an API key, a Claude subscription or Claude Platform on AWS, and your cloud provider's infrastructure on Amazon Bedrock or Google Cloud's Agent Platform. Through a gateway, the cache is wherever your requests are forwarded.

Can another organisation read text from my prompt cache?

No. Anthropic's API reference, read on 4 October 2026, says prompt caches are isolated between organisations, and that different organisations never share caches even when they use identical prompts. On the Claude API, Claude Platform on AWS and Microsoft Foundry, caches are also isolated between the workspaces inside one organisation. The same reference says the cached representations are held in memory only.

Does prompt caching change the answers Claude gives?

No. Prompt caching changes what the input costs, and the answer stays the same. Anthropic's API reference, read on 4 October 2026, says prompt caching has no effect on output token generation, and that the response is identical to the one you would get if prompt caching were not used. Output tokens have their own list price, for example $10 per million tokens on Claude Sonnet 5.5.

Is there a minimum size for a cached prompt?

Yes. The start of a request has to reach a minimum number of tokens before the prompt cache stores it. Anthropic's API reference, read on 4 October 2026, gives 512 tokens for a group of models that includes Claude Opus 5.5 and Claude Sonnet 5.5, 1,024 tokens for a group that includes Claude Sonnet 5, and 4,096 tokens for Claude Haiku 4.5. A shorter start is processed without caching, and the API returns no error.

Does Claude Code cache each file separately?

No. The prompt cache in Claude Code stores the start of the whole request as one piece. Anthropic's documentation, read on 4 October 2026, says: "There is no per-file or per-segment caching." A file that Claude read is part of the messages, so the service stores it together with everything that came before it in the request. The same page says a change anywhere in the start recomputes everything after it.

Can I clear the prompt cache myself?

No. Anthropic's API reference, read on 4 October 2026, says there is currently no way to clear the prompt cache yourself. Stored prefixes expire automatically after a minimum of 5 minutes without use, and each use of a stored copy starts the 5 minutes again at no extra charge. In Claude Code you can turn caching off by setting the `DISABLE_PROMPT_CACHING` environment variable to `1`.