The three layers of a Claude Code request and what changes each
A Claude Code request has three layers, in this order: the system prompt, the project context and the conversation. Anthropic's documentation, read on 4 October 2026, says Claude Code puts the content that rarely changes first, because the prompt cache matches the start of a request exactly. The system prompt changes when the set of loaded tool definitions changes. The project context changes when a session starts and after /clear or /compact. The conversation changes on every turn. A change in an early layer recomputes everything after it, so a change to the system prompt recomputes the whole request. This page explains each layer and where one stored copy can be reused.
Published October 4, 2026. Editorial.
Key takeaways
- Anthropic's documentation, read on 4 October 2026, orders a Claude Code request as system prompt, then project context, then conversation, with the content that rarely changes first.
- The system prompt layer holds the core instructions and the tool definitions, and Anthropic says a change to it invalidates everything after it.
- The project context layer holds CLAUDE.md, auto memory and unscoped rules, and it changes when a session starts or after /clear or /compact.
- Plan mode and skills add their instructions as conversation messages, so the stored start of the request stays intact, by Anthropic's own example.
- Anthropic says the cache in Claude Code is in effect limited to one machine and one directory, because the request contains the working directory and the auto memory paths.
Claude Code is Anthropic's coding tool: you type a request in a terminal, and an AI model reads files, runs commands and edits code. Every message you send becomes a new request to the model, and each request contains the whole conversation again. The work is counted in tokens, the pieces of text that the model processes.
The prompt cache is a store of request text that the service has already processed. When the start of a new request matches the stored text exactly, the service reuses that part at a lower price. Anthropic's name for the start of a request is the prefix. The page on what the prompt cache is covers the mechanism and the prices.
This page is about the order of a request. Claude Code builds each request in three parts, which Anthropic calls layers, and the order decides how much of the cache a change costs you. It belongs to the guide to the Claude Code prompt cache.
The three parts of a request, in order
Anthropic's documentation, read on 4 October 2026, says: "To get the most out of prefix matching, Claude Code orders each request so content that rarely changes between turns comes first." [1] A turn is one message from you plus the work Claude does in reply.
The first three columns of this table are Anthropic's own layer table [1]. The fourth column is our reading of what a change in that part costs, taken from the rules quoted in the sections below.
| Part of the request | What it holds | When it changes, in Anthropic's table | What a change costs |
|---|---|---|---|
| 1. System prompt | Core instructions, tool definitions | The set of loaded tool definitions changes | The whole request is recomputed. Anthropic's words: a change here "invalidates everything" |
| 2. Project context | CLAUDE.md, auto memory, unscoped rules | A session starts, or after /clear or /compact |
The project context and the conversation after it are recomputed. The system prompt before it still matches |
| 3. Conversation | Your messages, Claude's responses, tool results | Every turn | Text added at the end is the only new part. A change to an earlier message recomputes the conversation from that message onward |
Anthropic adds a caution about its third column: it gives common triggers and is shorter than the full set of causes [1]. The page on the nine actions that reset the cache lists the full set.
Part one: the system prompt
The system prompt is the set of instructions that Anthropic writes for the model. In the layer table it holds two things: the core instructions and the tool definitions [1]. A tool is an action the model can ask Claude Code to perform, such as reading a file or running a command. A tool definition is the description that tells the model what the tool does and how to call it.
Anthropic's interactive simulation of a session, read on 4 October 2026, shows a system prompt of 4,200 tokens and describes it as "Core instructions for behavior, tool use, and response formatting. Always loaded first. You never see it." [2] Anthropic calls the numbers in that simulation representative token counts, so your own session will differ [2].
The system prompt changes when the set of loaded tool definitions changes [1]. Two kinds of tool matter here.
The first kind is the tools from MCP servers. An MCP server is a program that connects Claude Code to an outside tool or data source. By default, on supported models, a feature named tool search holds back the full definitions of MCP tools until Claude needs them. In that case, Anthropic says, "Claude Code keeps the tool list from the conversation's first request for the whole conversation", so a server that connects or disconnects during the session does not disturb anything already stored [1]. When the tools load in full at the start instead, adding a definition invalidates the cache, which means the stored copy can no longer be used for the next request [1].
The second kind is the built-in tools. A deny rule is a permission rule that blocks Claude from using a tool. Anthropic's page says that if you add a bare tool name such as Bash as a deny rule while tool search is unavailable or disabled, Claude Code removes that definition from the next request, and the cache is invalidated [1].
The system prompt also contains one item that is specific to your computer. Anthropic's section on cache scope says: "The system prompt embeds your auto memory paths" [1]. Auto memory is the set of notes that Claude writes for itself about your project, and the paths are the places on your disk where those notes are kept.
Part two: the project context
The project context is the text Claude Code loads from your project and your own settings at the start of a session. Anthropic's table names three items [1]:
- CLAUDE.md is a file of instructions that you write and that Claude reads at the start of every session. Claude Code loads it from the folder you start in and from every folder above that one, and joins the files together [3].
- Auto memory is Claude's own notes. The first 200 lines of its index file,
MEMORY.md, or the first 25KB, whichever comes first, load at the start of every conversation [3]. - Unscoped rules are instruction files in the
.claude/rules/folder that have nopathsfield. Anthropic's memory page says these "are loaded at launch with the same priority as.claude/CLAUDE.md" [3].
This part comes after the system prompt. Anthropic's memory page says that CLAUDE.md content "is delivered as a user message after the system prompt" [3]. Our reading of that placement: by the exact-match rule, a different CLAUDE.md leaves the system prompt before it unchanged.
The project context changes at three moments in Anthropic's table: when a session starts, after /clear (the command that starts a new conversation) and after /compact (the command that replaces the conversation history with a summary) [1]. Between those moments it stays fixed. The same page says that project-root and user-level CLAUDE.md files "are read once at session start and held in memory", so an edit in the middle of a session waits until the next /clear, /compact or restart [1]. The page on actions that keep the cache covers that case in full.
In the simulation, the project CLAUDE.md is 1,800 tokens, the user CLAUDE.md is 320 tokens and auto memory is 680 tokens [2]. Added together that is 2,800 tokens. The sum is our own arithmetic on Anthropic's representative counts.
Part three: the conversation
The conversation holds your messages, Claude's responses and the tool results [1]. It changes on every turn, and the change is an addition at the end.
Several things that look like settings are in fact added to the conversation. Each of these statements is from Anthropic's page, read on 4 October 2026 [1]:
- File contents. "File contents enter context only when Claude reads them, and reads append to the conversation." If you edit a file that Claude read earlier, the earlier copy in the history stays as it was. Claude Code adds a notice that the file changed, and Claude reads it again if needed.
- Nested CLAUDE.md files and path rules. A CLAUDE.md file in a subfolder, or a rule file with a
pathsfield, loads when Claude first reads a matching file. After it loads, "the content is part of the conversation history". - Skills and commands. A skill is a packaged set of instructions for one task. Skills and commands "inject their instructions as user messages at the point of invocation. Nothing earlier in the conversation changes."
- Plan mode. Plan mode is a setting in which Claude researches and proposes changes without making them [5]. Its instructions are also added as conversation messages.
- Output style. When you change the output style in the middle of a session, Claude Code delivers the new style's instructions as a message in the conversation. Before v2.1.251, a style change in the middle of a session kept the cache but did not apply until you ran
/clearor started a new session.
The start of the conversation also holds facts about your computer. Anthropic's page says the conversation "opens with an announcement of the working directory, platform, shell, and OS version", and that each conversation also contains the branch and recent commits from a git status snapshot taken at startup [1]. Git is the tool that records the versions of your code, and the snapshot is a short record of its state at that moment.
Why an early change recomputes everything after it
The rule is one sentence in Anthropic's documentation: "The match is exact, so a change anywhere in the prefix recomputes everything after it. There is no per-file or per-segment caching." [1]
The API reference, read on 4 October 2026, describes the same rule as a hierarchy of tools, then system, then messages: "Changes at each level invalidate that level and all subsequent levels." [4]
The Claude Code page applies the rule to the three parts: "A change to the conversation layer leaves the system prompt and project context cached." A change to the system prompt, it continues, invalidates everything, because all the later content now comes after a different start [1]. The API reference explains why. A stored copy is identified by a calculation over all the text from the first character up to a marked point, so changing any text at or before that point produces a different result on the next request [4]. If the text near the start is different, each later point has a different history before it, and no stored copy matches.
The order of the three parts is Anthropic's design choice: the content that rarely changes comes first [1]. The system prompt is first. The project context changes at a few named moments, and it is second. The conversation changes on every turn, and it is last, so its changes leave the two earlier parts stored.
Two settings outside the table: model and effort level
Anthropic lists two settings that "don't appear in the layer table but still affect what stays cached" [1].
- Model. Each model has its own cache. Switching models recomputes the entire request even when the content is identical [1].
- Effort level. The effort level is the setting for how much the model reasons before it replies. On most models each effort level has its own cache, so a change in the middle of a session recomputes the entire request. On Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription, the cache stays intact by default [1].
The page on switching model or effort level in the middle of a task gives the details and a worked example.
Why plan mode and skills add to the end
Anthropic uses these two features as its examples of the exact-match rule: plan mode and skill loading "append their instructions as conversation messages, so the cached prefix stays intact" [1].
Our reading of the design: the alternative would be to put those instructions in the system prompt. By the rule above, that would change the first part of the request and recompute everything after it each time you entered plan mode or used a skill. Adding the instructions at the end costs only the new text.
There is one exception for plan mode. With the opusplan model setting, entering or leaving plan mode changes the model between Opus and Sonnet, and Anthropic's page says that this makes the mode change a model switch [1].
Where a stored copy can be reused: one computer and one folder
Anthropic's page says that in Claude Code "the cache is effectively scoped to one machine and directory" [1]. The reason is the content described above: the system prompt contains your auto memory paths, and the conversation opens with your working directory, platform, shell and operating system version. The page gives three cases [1]:
- Two sessions in different folders build different prefixes and miss each other's cache.
- Sessions you run at the same time in the same folder build matching prefixes and read each other's cache.
- Sessions you run one after another in the same folder share the prefix only when the git status snapshot taken at startup matches.
The API itself can share a stored copy more widely than this. Anthropic's page says caches are isolated between organisations and, on some providers, between workspaces inside an organisation. Inside those limits, "any two requests with the same model and prefix read the same cache" [1].
What to do with this
Our position: decide the first two parts of the request before you start, and make your changes in the third part. In practice that means three habits, each one a direct use of Anthropic's layer table.
- Set up your tools before the session. The tool definitions are in the first part, so a change there costs the whole request.
- Edit CLAUDE.md between sessions or between tasks. The project context loads at session start and after
/clearor/compact. - Choose the model and the effort level before the first message. Anthropic's tip on the same page says to choose both at the start of a session [1].
Reveneau is an AI software development consultancy. All of its code is written by AI, so token use is a running cost of every Reveneau build, and the order of a request is part of that cost. Reveneau is independent of Anthropic. The statements on this page are Anthropic's own descriptions of its product, read on 4 October 2026, and the fourth column of the table is our reading of them.
For the wider guide on using fewer tokens, see how to reduce Claude Code token usage.
Common questions
What are the three layers of a Claude Code request?
The three layers of a Claude Code request are the system prompt, the project context and the conversation, in that order. Anthropic's documentation, read on 4 October 2026, says Claude Code orders each request so that content which rarely changes between turns comes first. The system prompt holds core instructions and tool definitions, the project context holds CLAUDE.md, auto memory and unscoped rules, and the conversation holds messages, responses and tool results.
What is in the system prompt layer of a Claude Code request?
The system prompt layer of a Claude Code request holds the core instructions that Anthropic writes for the model and the tool definitions. Anthropic's documentation, read on 4 October 2026, also says the system prompt embeds your auto memory paths. In Anthropic's simulation of a session the system prompt is 4,200 tokens, a figure Anthropic calls a representative token count. The layer changes when the set of loaded tool definitions changes.
What is the project context layer in Claude Code?
The project context layer in Claude Code is the text loaded from your project at the start of a session: CLAUDE.md, auto memory and unscoped rules. Anthropic's layer table, read on 4 October 2026, says it changes when a session starts, or after `/clear` or `/compact`. In Anthropic's simulation, the project CLAUDE.md is 1,800 tokens, the user CLAUDE.md is 320 tokens and auto memory is 680 tokens.
Why does a change at the start of a request recompute the whole request?
A change at the start of a request recomputes the whole request because the prompt cache matches the start exactly. Anthropic's documentation, read on 4 October 2026, says a change anywhere in the prefix recomputes everything after it, and that there is no per-file or per-segment caching. The API reference describes a hierarchy of tools, system and messages, in which a change at one level invalidates that level and all later levels.
Is CLAUDE.md part of the system prompt?
No. CLAUDE.md is part of the project context layer, which comes after the system prompt. Anthropic's memory page, read on 4 October 2026, says CLAUDE.md content is delivered as a user message after the system prompt. Project-root and user-level CLAUDE.md files are read once at session start and held in memory, so an edit during a session waits until the next `/clear`, `/compact` or restart.
Why does plan mode keep the stored start of a request?
Plan mode keeps the stored start of a request because Claude Code adds the plan mode instructions as conversation messages at the end. Anthropic's documentation, read on 4 October 2026, uses plan mode as its example of the exact-match rule and says the cached prefix stays intact. The exception is the `opusplan` model setting, where entering or leaving plan mode changes the model between Opus and Sonnet, which Anthropic counts as a model switch.
Does loading a skill change the system prompt?
No. Loading a skill adds text to the conversation layer and leaves the system prompt as it was. Anthropic's documentation, read on 4 October 2026, says skills and commands inject their instructions as user messages at the point where you call them, and that nothing earlier in the conversation changes. The cost of the turn is the new skill text, and the stored start of the request still matches.
Do two Claude Code sessions share one prompt cache?
Two Claude Code sessions share a prompt cache when they build the same start of the request. Anthropic's documentation, read on 4 October 2026, says sessions you run at the same time in the same directory build matching prefixes and read each other's cache. Sessions in different directories build different prefixes and miss each other's cache, because the conversation opens with the working directory and the system prompt embeds the auto memory paths.
Why does a new session miss the cache of the session I ran before it?
A new session misses the cache of an earlier session in the same folder when the git status snapshot is different. Anthropic's documentation, read on 4 October 2026, says sessions run one after another share the prefix only when the snapshot taken at startup matches, because each conversation contains the branch and recent commits from that snapshot. The snapshot is a short record of the state of your code's version history at that moment.
Which layer do file contents belong to in a Claude Code request?
File contents belong to the conversation layer of a Claude Code request. Anthropic's documentation, read on 4 October 2026, says file contents enter context only when Claude reads them, and that reads are added to the end of the conversation. If you edit a file Claude read earlier, the earlier copy in the history stays as it was, and Claude Code adds a notice that the file changed.
Are the model and the effort level part of the three layers?
No. Anthropic lists the model and the effort level as two settings outside the layer table that still affect what stays cached. Its documentation, read on 4 October 2026, says each model has its own cache, so switching models recomputes the entire request even when the content is identical. On most models each effort level also has its own cache. On Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription, the cache stays intact by default.
Do nested CLAUDE.md files belong to the project context layer or the conversation layer?
Nested CLAUDE.md files belong to the conversation layer of a Claude Code request. Anthropic's documentation, read on 4 October 2026, says a CLAUDE.md file in a subfolder, or a rule file with a `paths` field, loads when Claude first reads a matching file, and that after it loads its content is part of the conversation history. Rules without a `paths` field are in the project context layer: Anthropic's memory page says they are loaded at launch with the same priority as `.claude/CLAUDE.md`.
References
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Explore the context window (code.claude.com), read 4 October 2026
- Anthropic, How Claude remembers your project (code.claude.com), read 4 October 2026
- Anthropic, Prompt caching (platform.claude.com), read 4 October 2026
- Anthropic, Choose a permission mode (code.claude.com), read 4 October 2026