Start here

How to read /usage and /context in Claude Code

/usage and /context are the two measuring commands in Claude Code. /usage shows what the session has used so far: token counts for each model, an estimated cost that Claude Code computes on your computer at list price, and, on a paid plan, a breakdown of what counts against your limits. /context shows what is in the context window at this moment, by category. Anthropic's documentation, read on 4 October 2026, says the cost figure is an estimate and names the Usage page in the Claude Console as the billing record. This page explains each line, adds the status line for a continuous reading, and gives a five-minute routine.

Published October 4, 2026. Editorial.

Key takeaways

  • In Claude Code, `/usage` reports a total that grows through the session, and `/context` reports the text that the next request will contain.
  • The Total cost line in `/usage` is an estimate computed on your computer at list price, and Anthropic's documentation names the Usage page in the Claude Console as the billing record.
  • In the sample that Anthropic's documentation prints, 940,000 of 996,500 tokens are cache reads, which is 94.3 percent by our arithmetic.
  • The plan usage breakdown in `/usage` flags a behaviour such as long context or cache misses when it accounts for 10 percent or more of recent usage.
  • The Claude Code status line runs on your computer and uses no tokens, and its `used_percentage` value counts input tokens only.

Claude Code, Anthropic's coding tool, has two commands that work as measuring tools. /usage reports how many tokens the session has used so far. /context reports what is in the context window at this moment. A token is a piece of text that the model processes, and the context window is the largest amount of text the model can read in one request.

The two commands answer different questions, and reading the wrong one leads to the wrong fix. This page explains each line of both screens, using Anthropic's documentation as read on 4 October 2026. It belongs to the guide on how to reduce Claude Code token usage. Our advice is short: measure before you change a setting, and measure again after.

Which screen answers which question

Your question Where to look What you get
How many tokens has this session used, and what would they cost? /usage, the Session block Token counts for each model and an estimated cost at list price [1]
Is the prompt cache working? /usage, the Prompt cache (main) line Share of input read from the cache, and the number of misses [1]
What counts against my plan limit? /usage, the plan usage breakdown Shares for skills, subagents, plugins and MCP servers [1]
Which repeating task uses the most? /usage, the Loops rows Total tokens and tokens per run for each scheduled task [1]
What is filling the context window now? /context A grid of current usage by category, with suggestions [3]
How full is the window, without asking? The status line A percentage that updates after each response [2]
How do I work, and where do requests go wrong? /insights A report on your recent sessions [1]
What will the organisation be charged? The Usage page in the Claude Console The billing record, which Anthropic calls authoritative [1]

The Session block of /usage, line by line

Type /usage in a session. Anthropic's command reference, read on 4 October 2026, describes it as showing "session cost, plan usage limits, and activity stats", and lists /cost and /stats as other names for the same command [3]. The Session block is at the top. This is the sample that Anthropic's documentation prints [1]:

Total cost:            $0.55
Total duration (API):  6m 20s
Total duration (wall): 6h 33m 10s
Total code changes:    0 lines added, 0 lines removed
Usage by model:
   claude-sonnet-4-6:  1.2k input, 5.3k output, 940.0k cache read, 50.0k cache write ($0.55)

Total cost is an estimate that Claude Code works out on your own computer. The documentation says Claude Code "computes the dollar figure locally from token counts at list price" and adds: "The figure is an estimate, so for authoritative billing see the Usage page in the Claude Console" [1]. The Claude Console is Anthropic's website for accounts that pay per token.

Two groups of people should read this line with care. People on a Pro or Max subscription have usage included in the subscription, and the documentation says the session cost figure "isn't relevant for billing purposes" for them [1]. Organisations with negotiated prices see list price here unless an administrator has set a table of the organisation's prices in a setting named modelPricing. When that table is in effect, the line shows the note at your organization's configured rates. Anthropic's documentation says this setting requires Claude Code v2.1.242 or later [1].

Total duration (API) is the time spent waiting for the model's responses. Total duration (wall) is the time the session has been running [2]. In the sample the first is 6 minutes 20 seconds and the second is 6 hours 33 minutes, so this session was open for hours while the model worked for minutes.

Total code changes counts the lines of code added and removed.

Usage by model is the most useful line. It gives four counts for each model:

  • input: new text the model read at full price
  • output: text the model wrote, which includes its reasoning
  • cache read: text read again from the prompt cache at a lower price
  • cache write: text stored in the prompt cache for the first time

The prompt cache is a store, kept by the service that runs the model, of request text it has already processed. Claude Code sends the whole conversation with every request, and the cache lets the service read the unchanged part again at a lower price.

What the sample line shows

It helps to read the sample closely, because it shows where the tokens of that one sample session went. The table below is our own arithmetic. It uses the list prices for Claude Sonnet 4.6 on Anthropic's pricing page, read on 4 October 2026: $3 per million input tokens, $15 per million output tokens, $0.30 per million cache read tokens and $3.75 per million tokens for a five-minute cache write [5].

Count in the sample Tokens Share of all tokens Cost at list price Share of the cost
Input 1,200 0.1 percent $0.0036 0.7 percent
Output 5,300 0.5 percent $0.0795 14.4 percent
Cache read 940,000 94.3 percent $0.2820 51.0 percent
Cache write 50,000 5.0 percent $0.1875 33.9 percent
Total 996,500 100 percent $0.5526 100 percent

The four costs add up to $0.5526, which rounds to the $0.55 in the sample. That match holds when the cache writes are priced at the five-minute rate. The sample does not say which rate applied, so this is our inference.

Three things follow from the table. First, 94.3 percent of the tokens are cache reads, which is the earlier conversation being read again on each request. Second, the text that was new to the model at full price is 1,200 tokens, which is one token in every 830 (996,500 divided by 1,200). Third, the output is half of one percent of the tokens and 14.4 percent of the cost, because output has the highest price.

In a session with these proportions, the way to lower the total is to make the conversation shorter, because new input is 0.1 percent of the tokens. The page on where Claude Code tokens go explains what the conversation contains.

When the totals reset

The documentation says these totals "reset when /clear starts a new session", so the next session's total cost starts at $0. Before Claude Code v2.1.211, the totals kept adding up across /clear for as long as the program stayed open [1]. If you compare two readings, check that no /clear came between them.

The prompt cache line

After the first response, the Session block also shows a line that begins Prompt cache (main). It gives the number of requests, the share of input tokens read from the cache, the number of cache misses, and whether the cache is still usable at this moment. Anthropic's documentation says the line requires Claude Code v2.1.251 or later and covers the main conversation only [1]. The page on how to check your cache hit rate explains each part of that line.

The plan usage breakdown

On a Pro, Max, Team or Enterprise plan, /usage also shows what counts against your plan limits. By Anthropic's documentation, read on 4 October 2026, the breakdown has three parts [1].

Attribution. Recent usage is divided between skills, subagents, plugins and individual MCP servers, each shown as a percentage of the total. A skill is a packaged set of instructions that loads when a task needs it. A subagent is a second copy of Claude that does one task in its own context window. A plugin is an installed package of add-ons. An MCP server is a program that connects Claude Code to an outside tool or data source. The documentation notes that an MCP server's share counts only the requests that used one of its tool results. Before v2.1.222, one call to a server caused every later request to be counted against it, which overstated its share [1].

Behaviour flags. Claude Code names a behaviour, "such as long context or cache misses", when it accounts for 10 percent or more of recent usage, and gives a tip to reduce it [1]. For a long context flag, use the commands on the page about /clear, /compact and /rewind.

Loops. A scheduled task is a prompt that Claude runs again on an interval, for example one started with the /loop command. The breakdown gives a row to each of the scheduled tasks that used the most tokens recently, ordered by total tokens. Each row reports how often the task runs, how many times it ran, its total tokens, its tokens per run, and when it last ran. Claude Code identifies a row by the task's prompt, so a loop that you stop and create again stays in one row. This part requires Claude Code v2.1.242 or later [1].

Press d or w to change between the last 24 hours and the last 7 days. One limit applies to the whole breakdown. The documentation says the figures "are approximate and computed from local session history on this machine", so work done on another device or on claude.ai is left out [1].

If the plan limits fail to load, most often because the usage service is limiting requests, /usage shows the last bars it loaded on this machine within the past 60 minutes, with a note that says how old they are. Press r to try again [1].

/context shows what is in the window now

/usage is a total that grows through the session. /context is a reading of one moment: the text that the next request will contain. Anthropic's command reference, read on 4 October 2026, says it shows current context usage as a coloured grid, with suggestions about tools that use a large share of the context, about memory files that have grown too large, and about a window that is close to full [3].

The breakdown is by category, and Anthropic's documentation says it includes "which CLAUDE.md and auto memory files loaded" [4]. CLAUDE.md is a file of instructions that you write and that Claude reads at the start of every session. Auto memory is a set of notes that Claude writes for itself. Both load before you type, so /context in a new session shows the fixed size that every request will contain.

Two details from the command reference, read on 4 October 2026 [3]:

  • When the conversation is larger than the context window, the output includes a warning that says how far over the limit you are and which command frees space.
  • In the fullscreen display mode, /context hides the list of single items so that the grid stays visible. Type /context all to show the list.

Use /context at three moments: at the start of a session, to see the fixed size; after a large task, to see what the task added; and before you start a long new task. Anthropic's advice on MCP servers is the same: "Run /context to see what's consuming space." [1] If MCP tools are a large category, read MCP servers or command-line tools.

The status line gives a continuous reading

Both commands need you to stop and ask. The status line shows a reading all the time. It is a bar at the bottom of Claude Code that runs a small script you choose and displays what the script prints. Claude Code can write the script for you. The documentation gives this example [2]:

/statusline show model name and context percentage with a progress bar

Claude Code passes the script a set of named values after each response. By Anthropic's documentation, read on 4 October 2026, four of them matter for token use [2]:

  • context_window.used_percentage: the share of the context window in use. It is calculated from input tokens only, which here means new input plus cache writes plus cache reads. Output tokens are left out.
  • context_window.context_window_size: the size of the window, which is 200,000 tokens by default and 1,000,000 for models with extended context.
  • cost.total_cost_usd: the same estimated session cost that /usage shows, at list price, which "may differ from your actual bill".
  • exceeds_200k_tokens: whether the most recent response passed 200,000 tokens in total. This threshold is fixed, whatever the size of the window.

The documentation states that the status line "runs locally and does not consume API tokens" [2]. In plain words, displaying it uses no tokens. Our recommendation is to show the context percentage and the session cost in it at all times.

/insights reports on how you work

/insights answers a different question. The documentation describes it as a report on how you work, where /usage reports how many tokens you used [1]. It reads your recent sessions on this machine and writes a web page that covers what you work on, where requests were misunderstood or code had faults, and suggestions.

Four facts from Anthropic's documentation, read on 4 October 2026 [1]:

  • One run analyses up to 200 sessions that it has not seen before.
  • The latest report is saved at ~/.claude/usage-data/report.html, and each run also keeps a dated copy.
  • The analysis itself uses tokens, and those tokens are counted against your plan or your pay-per-token account.
  • Sessions from other devices and from claude.ai are left out.

Reveneau recommends running it once a month. A misunderstood request uses tokens twice: once for the wrong work and once for the correction.

A measuring routine that takes five minutes

Reveneau is an AI software development consultancy in which all code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. This is the routine Reveneau recommends:

  1. At the start of a session, run /context and note the fixed size.
  2. Keep the context percentage in the status line.
  3. After each task, run /usage and read the Usage by model line. If cache reads are large, the conversation is long, and the next unrelated task should start after /clear.
  4. Once a week, run /usage, press w, and read the behaviour flags and the attribution shares.

When a number keeps rising and you cannot see why, the page on why usage keeps rising in a long session lists the eight causes that Anthropic documents.

Common questions

What does /usage show in Claude Code?

The `/usage` command in Claude Code shows session cost, plan usage limits and activity statistics. Its Session block lists an estimated total cost, two durations, the lines of code changed, and token counts for each model. On a Pro, Max, Team or Enterprise plan it also shows a breakdown of what counts against your plan limits. Anthropic's command reference, read on 4 October 2026, lists `/cost` and `/stats` as other names for it.

Is the Total cost in /usage what I will be billed?

No. The Total cost in `/usage` is an estimate that Claude Code computes on your computer from token counts at list price. Anthropic's documentation, read on 4 October 2026, says the Usage page in the Claude Console is the place for authoritative billing. For Pro and Max subscribers, usage is included in the subscription, so the documentation says this figure is irrelevant for their billing.

How do I read the Usage by model line in /usage?

The Usage by model line in `/usage` gives four token counts for each model: input, output, cache read and cache write. Input is new text read at full price, output is text the model wrote, and the two cache counts are text stored in or read from the prompt cache. In Anthropic's printed sample, read on 4 October 2026, the line shows 1.2k input, 5.3k output, 940.0k cache read and 50.0k cache write.

What is the difference between /usage and /context?

`/usage` is a total that grows through the session, and `/context` is a reading of one moment. `/usage` adds up every token the session has used since it started or since the last `/clear`. `/context` shows the text now in the context window, which is what the next request will contain. Anthropic's command reference, read on 4 October 2026, describes `/context` as a coloured grid of current context usage with suggestions.

What does /context show in Claude Code?

The `/context` command in Claude Code shows current context usage as a coloured grid, divided by category. Anthropic's documentation, read on 4 October 2026, says the breakdown includes which CLAUDE.md and auto memory files loaded, and that the command gives suggestions for reducing usage. When the conversation is larger than the context window, the output adds a warning that says how far over the limit you are.

Does /usage reset when I run /clear?

Yes. The `/usage` session totals reset when `/clear` starts a new session, so the next session's total cost starts at $0. Anthropic's documentation, read on 4 October 2026, says that before Claude Code v2.1.211 the totals kept adding up across `/clear` for as long as the program stayed open. The prompt cache line in the Session block resets at the same time.

What are the behaviour flags in /usage?

The behaviour flags in `/usage` are notices that name a pattern of use, such as long context or cache misses, when that pattern accounts for 10 percent or more of your recent usage. Each flag comes with a tip to reduce it. Anthropic's documentation, read on 4 October 2026, says the flags appear on Pro, Max, Team and Enterprise plans and are computed from session history on the same machine.

What do the Loops rows in /usage mean?

The Loops rows in `/usage` list the scheduled tasks that used the most tokens recently, one row for each task. A scheduled task is a prompt that runs again on an interval, such as one started with `/loop`. Each row gives how often the task runs, how many times it ran, its total tokens and its tokens per run. Anthropic's documentation, read on 4 October 2026, says the rows require Claude Code v2.1.242 or later.

Does the Claude Code status line use tokens?

No. The Claude Code status line runs a script on your own computer and uses no tokens, by Anthropic's documentation read on 4 October 2026. Claude Code passes the script values such as the context percentage and the estimated session cost after each response, and the script prints what you want to see. The command `/statusline`, followed by a description in plain words, asks Claude Code to write the script.

How is the context percentage in the status line calculated?

The context percentage in the status line is calculated from input tokens only: new input, plus tokens written to the cache, plus tokens read from the cache. Output tokens are left out. Anthropic's documentation, read on 4 October 2026, names the value `context_window.used_percentage` and gives the window size as 200,000 tokens by default, or 1,000,000 for models with extended context.

What does /insights do, and does it cost tokens?

The `/insights` command writes a report on how you work with Claude Code, and running it does use tokens. It analyses up to 200 recent sessions on your machine and covers what you work on, where requests were misunderstood, and suggestions. Anthropic's documentation, read on 4 October 2026, says the tokens it uses count against your plan or your pay-per-token account. The report is saved at `~/.claude/usage-data/report.html`.

Does /usage include usage from other devices?

No. The plan usage breakdown in `/usage` is computed from session history stored on the machine you are using. Anthropic's documentation, read on 4 October 2026, calls the figures approximate and says usage from other devices or from claude.ai is left out. Press `d` or `w` in the breakdown to change between the last 24 hours and the last 7 days on that machine.