Lifetime and checking

How to check your Claude Code cache hit rate

To check your Claude Code cache hit rate, run /usage and read the Prompt cache (main) line in the Session block. Anthropic's documentation, read on 4 October 2026, says the line shows the request count, the share of input tokens served from the cache, the number of misses, and whether the cache is warm, which means still inside its lifetime. It requires Claude Code v2.1.251 or later, and from v2.1.260 it also names a likely cause for the last miss when Claude Code can identify one. Anthropic's sample line shows 14 requests, 91 percent of input tokens from cache and 2 misses. A status line script can show the same numbers on every turn, and OpenTelemetry reports them for a whole organisation.

Published October 4, 2026. Editorial.

Key takeaways

  • The Prompt cache (main) line in /usage shows the request count, the share of input tokens from cache, the misses and whether the cache is still inside its lifetime, and it requires Claude Code v2.1.251 or later.
  • Anthropic's sample line, read on 4 October 2026, shows 14 requests, 91 percent of input tokens from cache, and 2 misses that wrote 310.2k tokens back to the cache.
  • Claude Code counts a request as a miss when it processed again more than 5 percent, and at least 2,000 tokens, of what it could have read from the cache.
  • From Claude Code v2.1.260 the line names a likely cause for the last miss, for example that the tool definitions changed.
  • The Prompt cache (main) line covers the main conversation only, and the OpenTelemetry metric claude_code.token.usage separates main, subagent and auxiliary requests.

Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. The work is counted in tokens, which are the pieces of text the model processes. The model keeps nothing between requests, so Claude Code sends the whole conversation again with every request.

The prompt cache is a store, kept by the service that runs the model, of request text it has already processed. When a new request begins with exactly the same text, the service reuses its earlier work and bills that text at a lower price. The cache hit rate is the share of your input tokens that were read from that store. This page shows the four places where Claude Code reports it, and how to trace a miss to its cause. It is part of the guide to the Claude Code prompt cache.

The two counts behind every figure

Every figure on this page is built from token counts that the API, the service that receives each request, reports on each response. Anthropic's documentation, read on 4 October 2026, names the two that matter for the cache [1]:

  • cache_creation_input_tokens: tokens written to the cache on this turn, billed at the cache write price. This is called a cache write.
  • cache_read_input_tokens: tokens served from the cache on this turn, billed at the cached token price, which is below the standard input price. This is called a cache read.

A third count, input_tokens, holds only the input that was neither read from the cache nor written to it. The API reference gives the sum: total input tokens are cache_read_input_tokens plus cache_creation_input_tokens plus input_tokens [5]. A small input_tokens number beside a large cache read number is therefore normal when the cache is working.

Anthropic's reading rule is short: "A high read-to-creation ratio means caching is working well. If creation stays high turn after turn, something is changing in your prefix." [1] The prefix is the start of the request, the part that the service compares with text it processed recently. A request that finds a stored copy of its start is a cache hit. A request that finds none is a cache miss.

The Prompt cache (main) line in /usage

The first place to look is the /usage command, which opens a screen with the token counts and cost of the current session. After the first response of the main conversation, Claude Code adds a line named Prompt cache (main) to the Session block of that screen. This requires Claude Code v2.1.251 or later [2]. Anthropic's page on costs, read on 4 October 2026, shows this sample [2]:

Prompt cache (main):   14 requests · 91% of input tokens from cache · 2 misses (last 6m 10s ago, 310.2k tokens re-cached) · 1 expected rebuild (compaction or tool-result clearing) · warm (1h TTL, last activity 40s ago)

Read it from left to right. Each part is defined on the same page [2].

Requests. The number of requests in the main conversation this session. The sample shows 14.

Share from cache. The share of input tokens that were served from the cache. The sample shows 91 percent. This is the cache hit rate.

Misses. Requests that processed again content the cache already held. The line gives the time of the last miss and how many tokens those requests wrote back to the cache. In the sample, 2 misses wrote 310.2k tokens.

Likely cause. When Claude Code can identify a likely cause for the last miss, the line names it, for example likely cause: tool definitions changed. This text requires Claude Code v2.1.260 or later.

Expected rebuilds. Claude Code sometimes rewrites the conversation itself. It does this by compaction, the step that replaces the conversation history with a summary, or by clearing old tool results from the conversation. A tool result is what comes back when Claude reads a file or runs a command. Claude Code counts the miss that follows as an expected rebuild and keeps it out of the miss count. This part of the line appears only after one has happened.

Warm or cold. Whether the stored copy is still inside its lifetime. The lifetime is called the time to live, or TTL: how long an unused stored copy is kept. A cache inside its lifetime is warm, and the line shows the TTL in effect. A cache whose lifetime has passed is cold, and the line then shows how long the session has been idle.

Three limits apply to this line, all stated by Anthropic [2]. It covers the main conversation only, so the requests of subagents (helpers that work in their own conversation) are left out. The /clear command, which starts a new conversation, resets it together with the Session block. And it is built from the cache token counts in the API's responses, so it works with every model provider and every gateway. A gateway is a server that an organisation runs between Claude Code and the model provider.

The related guide's page on how to read /usage and /context explains the other lines of the screen.

What each reading means and what to do

Anthropic gives no target number for the share from cache. Its only guidance is the ratio rule quoted above, and its sample shows 91 percent. The table turns each reading into an action. The meanings are Anthropic's [1] [2] [3]. The actions are our recommendation.

What you see What it means What to do
A high share from cache and no misses Most of the input is read from the store Nothing
2 misses (last 6m 10s ago, 310.2k tokens re-cached) Two requests processed again content the cache already held Read the likely cause, then compare what you did at that time with the list of actions that reset the cache
likely cause: tool definitions changed Claude Code identified a change in the set of tool definitions as the likely cause of the last miss Check whether an MCP server connected or a tool was removed during the session
1 expected rebuild Claude Code compacted the conversation or cleared old tool results Nothing to repair. Choose when compaction happens by running /compact between tasks
warm (1h TTL, ...) The stored copy is inside its lifetime Keep working
Cold, with an idle time The lifetime passed while the session was idle Expect the next request to store the conversation again. If this repeats, consider the one-hour lifetime
no prompt caching reported by the API No response in this session has reported cache tokens Check whether caching is disabled, and test your gateway or provider

An MCP server is a program that gives Claude extra tools. The page on the nine actions that reset the cache is the list to compare against. The page on how long the cache lasts covers the cold case.

A status line script that shows the cache on every turn

The status line is a bar at the bottom of Claude Code that runs a shell script you choose. Anthropic's documentation, read on 4 October 2026, says the script receives session data as JSON (a text format for structured data), and the bar displays whatever the script prints [3]. Anthropic calls a status line script the most direct way to watch the cache counts live [1].

Two objects in that data concern the cache [3]:

  • context_window.current_usage holds the token counts of the last API call, including cache_creation_input_tokens and cache_read_input_tokens. It is null before the first API call in a session, and again after /compact until the next API call.
  • prompt_cache holds the session statistics for the main conversation, the same numbers as the /usage line. It appears after the first API response of the main conversation and requires Claude Code v2.1.251 or later.

The fields of prompt_cache that a short script is likely to use are these, from Anthropic's field table [3]:

Field What it holds
warm Whether the stored start is still inside its TTL. It is also false when the last response reported no cache tokens
hit_ratio Cache read tokens as a fraction of all input tokens this session, from 0 to 1. The total counts cache reads, cache writes and uncached input
misses Requests that processed again content the cache already held
expected_rebuilds Rebuilds that followed a compaction or a clearing of old tool results
ttl The lifetime of the current stored start: "5m" or "1h"
caching_observed Whether any response this session reported cache tokens
miss_recache_tokens Tokens written to the cache by the requests counted as misses
recache_tokens_if_cold Tokens the next request stores again if the cache has expired by then

Anthropic also defines what counts as a miss. A request is a miss when it processed again more than 5 percent, and at least 2,000 tokens, of what it could have read from the cache, with no compaction or clearing of tool results to explain the difference [3].

The script below is our own example, built from the fields Anthropic documents, so test it before you rely on it. It uses jq, the command-line tool for reading JSON that Anthropic's own examples use [3].

#!/bin/bash
# Read the JSON that Claude Code sends on standard input
input=$(cat)

# Token counts from the last API response ("// 0" gives 0 when the value is null)
READ=$(echo "$input" | jq -r '.context_window.current_usage.cache_read_input_tokens // 0')
WRITE=$(echo "$input" | jq -r '.context_window.current_usage.cache_creation_input_tokens // 0')

# Session statistics. The prompt_cache object is absent until the first response.
STATE=$(echo "$input" | jq -r 'if .prompt_cache == null then "no data yet" elif .prompt_cache.warm then "warm" else "cold" end')
HIT=$(echo "$input" | jq -r '.prompt_cache.hit_ratio // 0')
MISSES=$(echo "$input" | jq -r '.prompt_cache.misses // 0')

echo "cache read $READ | cache write $WRITE | hit ratio $HIT | misses $MISSES | $STATE"

To use it, Anthropic's steps are: save the script to a file such as ~/.claude/statusline.sh, make it executable with chmod +x ~/.claude/statusline.sh, and add this to your user settings file at ~/.claude/settings.json [3]:

{
  "statusLine": {
    "type": "command",
    "command": "~/.claude/statusline.sh"
  }
}

Given the sample data on Anthropic's page, the script prints cache read 2000 | cache write 5000 | hit ratio 0.91 | misses 2 | warm. It prints cold whenever warm is false, which by Anthropic's field table includes the case where the last response reported no cache tokens [3]. Claude Code runs the script again when a new assistant message arrives, and also when a warm prompt cache in the data reaches its expires_at time [3].

The same sample holds a useful sum. It shows 352,000 tokens written to the cache in the session and 310,200 of them written by the two misses [3]. By our arithmetic, the misses account for 88.1 percent of the cache writes in that sample session.

The cause of the last miss, as Claude Code reports it

From Claude Code v2.1.260, the prompt_cache object has a field named last_miss_cause. Anthropic's documentation, read on 4 October 2026, says its causes list holds one or more cause names, and gives four examples [3]:

  • tools_changed: comes with counts of how many tools were added to or removed from the request.
  • system_prompt_changed: comes with the change in the length of the system prompt, in characters. The system prompt is the first part of every request: Anthropic's instructions for the model and the tool definitions.
  • ttl_expired_5m
  • likely_server_side

Anthropic lists the last two names without a description. The field is null until the first miss of the session, and whenever Claude Code could not identify a cause for the most recent miss [3]. A second field, miss_causes, counts how many of the session's diagnosed misses had each cause [3].

Checking a whole organisation with OpenTelemetry

OpenTelemetry is an open standard for sending measurements from a program to a monitoring system. Claude Code can export its usage data this way. Anthropic's documentation, read on 4 October 2026, says the export reports cache read tokens and cache creation tokens for each user and session [1].

You turn it on with environment variables, which are named values that a program reads from the shell when it starts. Anthropic's quick start sets CLAUDE_CODE_ENABLE_TELEMETRY=1 and then chooses an exporter, for example OTEL_METRICS_EXPORTER=otlp [4].

Two parts of the export hold cache counts [4]:

  • The metric claude_code.token.usage counts tokens and has a type attribute with four values: input, output, cacheRead and cacheCreation. The input type leaves out tokens read from or written to the cache. A query_source attribute separates main, subagent and auxiliary requests.
  • The event claude_code.api_request is logged for each request and has cache_read_tokens and cache_creation_tokens.

To get one hit rate for a team, apply Anthropic's definition of hit_ratio to the metric: divide the cacheRead total by the sum of the input, cacheRead and cacheCreation totals. That formula is our application of Anthropic's definition [3] [4]. Unlike the /usage line, the metric includes subagent requests, so filter by query_source when you want the main conversation alone.

How to trace a miss to its cause

Use the same order each time.

  1. Read the likely cause. If the /usage line names one, start there.
  2. Compare with your own actions. Anthropic's list of actions that reset the cache includes switching models, changing the effort level (the setting for how much the model reasons), connecting or removing an MCP server, and compacting [1]. Ask which one you did just before the time of the last miss.
  3. Check the lifetime. If the line says cold, the lifetime passed while the session was idle, and the next request stores the conversation again. Compare the length of your pauses with the lifetime in effect, five minutes or one hour.
  4. Watch cache writes over several turns. If cache writes stay high turn after turn, Anthropic says something is changing in the start of your request [1].
  5. Check the connection. If the line ends with no prompt caching reported by the API, no response has reported cache tokens. For the field caching_observed, Anthropic says a value of false means prompt caching is off, or your provider or gateway does not report it [3]. The page on gateways and cloud providers covers that case.

Turning caching off for a test

Anthropic's documentation, read on 4 October 2026, says that disabling caching is "occasionally useful when debugging caching behavior with a specific model or provider" [1]. You do it by setting one of these environment variables to 1 [1]:

Variable Effect
DISABLE_PROMPT_CACHING Disables caching for all models
DISABLE_PROMPT_CACHING_HAIKU Disables caching for the default Haiku model
DISABLE_PROMPT_CACHING_SONNET Disables caching for the default Sonnet model
DISABLE_PROMPT_CACHING_OPUS Disables caching for the default Opus model
DISABLE_PROMPT_CACHING_FABLE Disables caching for Fable only

Anthropic's list of environment variables adds that DISABLE_PROMPT_CACHING takes precedence over the per-model variables [6].

The per-model variables apply to the model that the short name resolves to. Anthropic's example: a session on claude-sonnet-5 keeps caching with DISABLE_PROMPT_CACHING_SONNET set while the name sonnet resolves to claude-sonnet-5-5. The Haiku variable covers the main conversation from Claude Code v2.1.283 [1].

Remove the variable when the test is over. Anthropic's sentence on this is plain: "For normal use, leave caching enabled." [1]

Our position

Check the Prompt cache (main) line at the end of a long task, before you start the next one. It takes one command, and it names the miss count, the tokens those misses stored again, and the likely cause when Claude Code can identify one. Reveneau recommends the status line script for anyone who works in sessions of several hours, because it shows a miss on the turn where it happens.

Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. Every figure on this page is Anthropic's own statement about its own product, and the sums marked as ours are arithmetic on Anthropic's sample data.

Common questions

How do I check my Claude Code cache hit rate?

To check your Claude Code cache hit rate, run `/usage` and read the `Prompt cache (main)` line in the Session block. Anthropic's documentation, read on 4 October 2026, says the line shows the share of input tokens served from the cache, the request count, the misses and whether the cache is still inside its lifetime. It appears after the first response of the main conversation and requires Claude Code v2.1.251 or later.

What is a good cache hit rate in Claude Code?

Anthropic gives no target number for a good cache hit rate in Claude Code. Its documentation, read on 4 October 2026, gives a rule in its place: a high ratio of cache reads to cache writes means caching is working well, and cache writes that stay high turn after turn mean something is changing in the start of your request. Anthropic's sample `/usage` line shows 91 percent of input tokens from cache.

What does the Prompt cache (main) line in /usage show?

The `Prompt cache (main)` line in `/usage` shows how the main conversation has used the prompt cache in this session. By Anthropic's documentation, read on 4 October 2026, it gives the request count, the share of input tokens from cache, the misses with the time of the last one and the tokens stored again, any expected rebuilds, and whether the cache is warm (inside its lifetime) or cold (past it), with the lifetime in effect.

What counts as a cache miss in Claude Code's statistics?

In Claude Code's statistics, a cache miss is a request that processed again content the cache already held. Anthropic's documentation, read on 4 October 2026, sets the threshold: more than 5 percent, and at least 2,000 tokens, of what the request could have read from the cache, with no compaction or clearing of tool results to explain it. A miss that follows compaction is counted separately as an expected rebuild.

What does likely cause: tool definitions changed mean in /usage?

The text `likely cause: tool definitions changed` in `/usage` means Claude Code identified a change in the set of tool definitions as the likely cause of the last cache miss. Anthropic's documentation, read on 4 October 2026, says this text requires Claude Code v2.1.260 or later. In the status line data, the matching cause name is `tools_changed`, and it comes with counts of the tools added and removed.

What does no prompt caching reported by the API mean?

The text `no prompt caching reported by the API` at the end of the `Prompt cache (main)` line means that no response in the session has reported cache tokens. That is Anthropic's documentation, read on 4 October 2026. For the matching status line field, `caching_observed`, Anthropic says a value of `false` means prompt caching is off, or your provider or gateway does not report it.

Which status line fields show the prompt cache?

Two objects in the status line data show the prompt cache. By Anthropic's documentation, read on 4 October 2026, `context_window.current_usage` holds the cache read and cache write token counts of the last API call, and `prompt_cache` holds the session statistics, including `warm`, `hit_ratio`, `misses` and `ttl`. The `prompt_cache` object appears after the first API response of the main conversation and requires Claude Code v2.1.251 or later.

Does the Prompt cache (main) line include subagent requests?

No. The `Prompt cache (main)` line covers the main conversation only. Anthropic's documentation, read on 4 October 2026, says it leaves out subagents, and that Claude Code does not count subagent requests in the `prompt_cache` statistics either. To see subagent requests, use the OpenTelemetry metric `claude_code.token.usage`, whose `query_source` attribute separates `main`, `subagent` and `auxiliary` requests.

How can an organisation measure cache reads for all its developers?

An organisation can measure cache reads for all its developers with OpenTelemetry, an open standard for sending measurements to a monitoring system. Anthropic's documentation, read on 4 October 2026, says the export reports cache read and cache creation tokens for each user and session. You enable it with `CLAUDE_CODE_ENABLE_TELEMETRY=1`. The metric `claude_code.token.usage` has a `type` attribute with the values `input`, `output`, `cacheRead` and `cacheCreation`.

Why is the input token count so small next to the cache read count?

The input token count is small because it leaves out everything that was read from or written to the cache. Anthropic's API reference, read on 4 October 2026, says `input_tokens` holds only the input that was neither read from the cache nor written to it. Total input is `cache_read_input_tokens` plus `cache_creation_input_tokens` plus `input_tokens`. A small input count beside a large cache read count is normal when the cache is working.

How do I turn off prompt caching in Claude Code for a test?

To turn off prompt caching in Claude Code for a test, set the environment variable `DISABLE_PROMPT_CACHING` to `1`. Anthropic's documentation, read on 4 October 2026, says this disables caching for all models, and it lists per-model variables for the default Haiku, Sonnet and Opus models and for Fable. Anthropic describes this as occasionally useful when debugging, and says to leave caching enabled for normal use.

Which Claude Code version shows prompt cache statistics?

Claude Code v2.1.251 or later shows prompt cache statistics. Anthropic's documentation, read on 4 October 2026, gives that version for both the `Prompt cache (main)` line in `/usage` and the `prompt_cache` object in the status line data. The likely cause of the last miss, and the fields `last_miss_cause` and `miss_causes`, require Claude Code v2.1.260 or later.