Choices

Which Claude Code model and effort level to use for which task

Use Sonnet at medium effort for most Claude Code work, Opus for architecture decisions, and Haiku for simple tasks given to a subagent. The choice of model is Anthropic's own advice, read on 4 October 2026, medium is the default effort on Sonnet 5.5, and the price list supports both: Claude Opus 5.5 is listed at $4 per million input tokens and $20 per million output tokens, twice the price of Claude Sonnet 5.5. The effort level sets how much the model reasons before it replies, and reasoning is billed as output. Opus 5.5 is the default model on Anthropic's plans, so the saving starts with one command, /model sonnet. This page gives the prices, the five effort levels and a table of tasks.

Published October 4, 2026. Editorial.

Key takeaways

  • Anthropic's documentation, read on 4 October 2026, advises Sonnet for most coding tasks, Opus for complex architecture decisions or multi-step reasoning, and `model: haiku` for simple subagent tasks.
  • The `default` model setting in Claude Code selects Opus 5.5 on Pro, Max, Team and Enterprise plans and on the Anthropic API (Anthropic's pay-per-token service), so a developer who chooses nothing works on Opus.
  • Claude Opus 5.5 is listed at $4 input and $20 output per million tokens against $2 and $10 for Claude Sonnet 5.5, and both list a cache read at $0.20, on Anthropic's pricing page read on 4 October 2026.
  • Opus 5.5 and Sonnet 5.5 start at `medium` effort, thinking cannot be turned off on them or on the Fable models, and a positive `MAX_THINKING_TOKENS` value has no effect on them.
  • Fast mode on Opus 5.5 is listed at $8 input and $40 output per million tokens, twice the standard Opus 5.5 price, and on a subscription it is paid from usage credits.

Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. Two settings affect what every request costs. The first is the model. The second is the effort level, which is the setting for how much the model reasons before it replies. This page is part of the guide on how to reduce Claude Code token usage, and it covers both settings.

All usage is counted in tokens. A token is a piece of text that the model processes, and Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters or 0.75 of an English word [3]. Each model has its own price per token, and the model's reasoning is billed at the highest of those prices. The model sets the price of every token. The effort level changes how many of the most expensive tokens the model writes.

Anthropic's advice in three lines

Anthropic's documentation for Claude Code, read on 4 October 2026, gives its advice in one short paragraph [1]:

  • Sonnet "handles most coding tasks well and costs less than Opus".
  • Opus is for "complex architectural decisions or multi-step reasoning".
  • For a simple task given to a subagent, set model: haiku. A subagent is a second copy of Claude that works on one task in its own separate conversation and sends back a summary.

Sonnet, Opus and Haiku are three families of models. A fourth family, Fable, is described in the same documentation as the most capable models in Claude Code, for tasks too large to finish in one working session [2].

One more fact makes this advice useful. The model you get when you choose nothing is Opus. The documentation says the default setting selects Opus 5.5 on Pro, Max, Team and Enterprise plans and on the Anthropic API, which is Anthropic's own pay-per-token service [2]. A developer who never opens the model menu therefore works on the model that Anthropic advises keeping for the hardest tasks. The same documentation names "Opus left as the default model" as one of the two usual causes of unexpectedly high spending on an API or cloud provider plan [1].

The model names and what each one selects

You choose a model in Claude Code with a short name called an alias. The table lists the aliases that matter for cost, with Anthropic's own description of each, read on 4 October 2026 [2].

Alias What Anthropic says it is for What it selects on the Anthropic API
sonnet "daily coding tasks" Sonnet 5.5
opus "complex reasoning tasks" Opus 5.5
haiku "simple tasks", on a "fast and efficient" model A Haiku model. The page does not name the version
fable "your hardest and longest-running tasks" Fable 5.1
opusplan Opus while you plan, Sonnet while Claude writes the code Both, one after the other
default Removes your own choice and returns to the default for your account Opus 5.5 on Anthropic's plans and API

Three details from the same page prevent surprises [2]:

  • The alias depends on where you buy the model. On Amazon Bedrock and Google Cloud's Agent Platform, sonnet selects Sonnet 4.5. On Microsoft Foundry, opus selects Opus 4.6 and the default is Sonnet 4.5.
  • New models need a recent version. Sonnet 5.5 requires Claude Code v2.1.284 or later, Opus 5.5 requires v2.1.280 or later, and Fable 5.1 requires v2.1.257 or later.
  • Fable can be billed separately. No plan uses a Fable model as its default. Depending on your plan, Fable usage can be billed to usage credits, which are paid usage beyond what your plan includes. In an interactive session, Claude Code asks for your consent before the first such request.

You switch with the command /model sonnet, or start a session with claude --model sonnet. Typing /model with a name also saves that model as your default for new sessions [2].

What each model costs per token

The prices below are Anthropic's list prices in US dollars per million tokens, from Anthropic's pricing page read on 4 October 2026 [3]. Haiku 4.5 is the only Haiku model on that page that is not marked as retired.

Model Input Cache read Output
Claude Haiku 4.5 $1 $0.10 $5
Claude Sonnet 5.5 $2 $0.20 $10
Claude Opus 5.5 $4 $0.20 $20
Claude Fable 5.1 $10 $0.25 $50

Three terms in that table need a plain explanation. Input is text the model reads. Output is text the model writes, which includes its reasoning. A cache read comes from the prompt cache, which is a store of request text that the service has already processed. Text read again from that store is billed at the lower cache read price. The related guide explains what the prompt cache is.

The ratios are our own arithmetic on Anthropic's figures. Opus 5.5 costs twice what Sonnet 5.5 costs for input and for output. Fable 5.1 costs five times what Sonnet 5.5 costs for both. Haiku 4.5 costs half. The cache read price is the same on Opus 5.5 and Sonnet 5.5, at $0.20.

Here is an invented example, with numbers chosen only to show the arithmetic. One request reads 100,000 tokens from the cache and writes 2,000 tokens of output.

Model Cache read cost Output cost Total for the invented request
Claude Haiku 4.5 $0.01 $0.01 $0.02
Claude Sonnet 5.5 $0.02 $0.02 $0.04
Claude Opus 5.5 $0.02 $0.04 $0.06
Claude Fable 5.1 $0.025 $0.10 $0.125

In this example the whole difference between Sonnet 5.5 and Opus 5.5 comes from the output. That is the reason the effort level matters as much as the model: reasoning is output.

Two limits on this comparison. First, the pricing page says that Claude 4.7 and later models use a newer tokenizer, the part that splits text into tokens, and Anthropic puts the increase at 30 percent more tokens for the same text [3]. The same file can therefore count as more tokens on a newer model. Second, on a Pro or Max subscription the usage is included in the plan, and the documentation says the dollar figure in /usage "isn't relevant for billing purposes" there [1]. The model still matters on a subscription, because usage counts against the plan's limits. The page on how to read /usage and /context shows where the usage for each model appears.

Effort level: how much the model reasons

Before it answers, the model can write out its reasoning. Anthropic calls this extended thinking. The documentation, read on 4 October 2026, says extended thinking is on by default, that "Thinking tokens are billed as output tokens", and that the default allowance "can be tens of thousands of tokens per request depending on the model" [1]. You pay for reasoning you never see: "You are charged for all thinking tokens generated, even when collapsed or redacted." [2]

The effort level controls this. On current models the model decides at each step whether to reason and for how long, and the effort level sets how much reasoning it tends to use [2]. Anthropic's table for choosing a level reads as follows [2].

Level When Anthropic says to use it
low Short exchanges where you review each result, such as a first draft or a small change like a rename
medium Everyday engineering work with a clear scope, such as a new feature. The default on Opus 5.5 and Sonnet 5.5
high Work where checking matters or unusual cases are likely, such as a bug fix in existing code. The default on every model except Opus 5.5, Sonnet 5.5 and Opus 4.7
xhigh Deeper reasoning at a higher token cost. The default on Opus 4.7
max Hard problems you want Claude to work through without you, such as a search for security faults

Anthropic adds a warning about max: it can add cost without adding quality, it tends to reason more than the task needs, and you should test it before using it widely [2].

Four facts from the same page, read on 4 October 2026, help you set it correctly [2]:

  • The levels differ by model. Fable 5.1, Fable 5, Opus 5.5, Sonnet 5.5, Opus 5, Sonnet 5, Opus 4.8 and Opus 4.7 offer all five levels. Opus 4.6 and Sonnet 4.6 offer four, with no xhigh. A model outside that list has no effort setting, and no Haiku model is on the list.
  • The same name means a different amount on each model. Anthropic states that in its own testing Opus 5.5 at medium matches or exceeds Opus 5 at high, and it advises starting Opus 5.5 at medium.
  • You set it with /effort. Type /effort low to set a level, /effort alone to open a slider, or /effort auto to remove your saved level. Pressing Enter saves the level for that model in later sessions. Pressing s applies it to the current session only, which requires Claude Code v2.1.257 or later. The max level always applies to the current session only, unless you set it through CLAUDE_CODE_EFFORT_LEVEL. That is an environment variable, a named setting that a program reads when it starts.
  • One word raises reasoning for one request. If you include ultrathink in a request, Claude Code asks for deeper reasoning on that turn, and the session's effort level stays the same.

Our recommendation follows from the last point. Keep the session at the default level, and add ultrathink to the one request that needs more reasoning. That costs less than raising the level for every request in the session.

Models where thinking stays on, and the fixed budget

Thinking can be turned off on some models and stays on in others. The documentation, read on 4 October 2026, says: "You can't turn off thinking on Opus 5.5, Sonnet 5.5, or the Fable models, which always use extended thinking." [1] On those models the switch in /config shows the text Thinking can't be turned off, and a saved setting of MAX_THINKING_TOKENS=0 has no effect [2].

On other models you have three controls. You can switch thinking for the current session with Option+T on macOS or Alt+T on Windows and Linux. You can set the default in /config. You can set MAX_THINKING_TOKENS=0, which turns thinking off on the Anthropic API [2].

MAX_THINKING_TOKENS has a second use: a positive number sets a fixed allowance of thinking tokens. This works only on models that use a fixed thinking budget. The documentation says Fable models, Sonnet 5 and later, and Opus 4.7 and later "always use adaptive reasoning", which means they choose their own amount and ignore a positive number [2]. On Opus 4.6 and Sonnet 4.6, a fixed allowance applies only after you set a second variable, which returns those two models to the fixed budget [2]. The two lines below combine the two settings the documentation names, with Anthropic's own example value of 8000 [1].

# Opus 4.6 and Sonnet 4.6: use a fixed thinking allowance of 8,000 tokens
export CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1
export MAX_THINKING_TOKENS=8000

On Opus 5.5 and Sonnet 5.5, which are what opus and sonnet select on the Anthropic API, the effort level is the main control over reasoning. Anthropic's documentation calls it the primary control [2].

Which model and effort level for which task

The pairings in this table are our recommendation. Each row combines Anthropic's advice on models [1] with Anthropic's table of effort levels [2], both read on 4 October 2026.

Task Model Effort level
A rename or another small change that you review step by step sonnet low
A new feature with a clear scope sonnet medium, the default on Sonnet 5.5
A bug fix in existing code, where unusual cases are likely sonnet high
An architecture decision, or reasoning across many steps opus medium, the default on Opus 5.5
A search, a test run or a log read that you give to a subagent haiku None. Haiku has no effort setting
A search for the original cause of a fault, or a long task left to run by itself fable high, the default on Fable
A hard problem that Claude works through without you opus or fable max, after you have tested it

Start each task on the lower-cost choice, and move up when the result is wrong. A wrong result on Sonnet costs one retry. Opus 5.5 on every request means twice the Sonnet 5.5 list price for new input and for output.

Subagents: put simple work on Haiku

A subagent runs on the main conversation's model unless something says otherwise. The documentation, read on 4 October 2026, gives the order Claude Code checks: a model that Claude passes for that one call, then the model field in the subagent's definition, then the CLAUDE_CODE_SUBAGENT_MODEL environment variable, then the main conversation's model [5]. It also states the consequence for cost: "A switch to Opus also applies to the subagents that inherit your session's model." [1]

Three settings move subagent work to a lower-cost model [5]:

  • Set model: haiku in the definition of a subagent you wrote.
  • Create a subagent named Explore with model: haiku. It replaces the built-in Explore subagent, which searches code and otherwise runs on the main conversation's model.
  • Set two environment variables to run every subagent on one model. This requires Claude Code v2.1.257 or later.

The documentation gives this settings block for the third option [5]:

{
  "env": {
    "CLAUDE_CODE_SUBAGENT_MODEL": "haiku",
    "CLAUDE_CODE_SUBAGENT_MODEL_FORCE": "1"
  }
}

To check the result, run /tasks while a subagent is working. From Claude Code v2.1.242, its row names the model [5]. A hook is a script that Claude Code runs by itself at a fixed point, and the page on hooks that trim test and log output compares a hook with a subagent for long output.

opusplan and fast mode

opusplan uses Opus in plan mode and Sonnet afterwards [2]. Plan mode is a setting in which Claude reads code and proposes a plan, and edits nothing until you approve. The pairing gives you Opus for the decision and the Sonnet price for writing the code. It has one cost that is hard to notice. Each model keeps its own prompt cache, so the documentation says each change into or out of plan mode is a model switch that starts a new cache [6]. Use opusplan when you plan once at the start of a task. Avoid it if you move in and out of plan mode many times in one session.

Fast mode makes Opus respond sooner and costs more per token. Anthropic's documentation, read on 4 October 2026, describes it as "up to 2.5x faster at a higher cost per token", says it is a research preview, meaning an early release that may change, and limits it to Opus 5.5, Opus 5 and Opus 4.8 [4]. Anthropic's pricing page lists fast mode on Opus 5.5 at $8 per million input tokens and $40 per million output tokens [3], which is twice the standard Opus 5.5 price by our arithmetic. On a subscription plan it is paid from usage credits and is outside the plan's included limits [4]. If you turn it on in the middle of a conversation, you pay the full fast mode input price once for the whole conversation so far, so Anthropic advises turning it on at the start [4]. Leave fast mode off unless your waiting time costs more than the tokens.

Choose at the start of the session

Pick the model before the first request and keep it. The documentation states the reason: "Each model has its own cache", so after a switch with /model the next request reads the whole conversation with no cache reads [6]. The page on switching model or effort in the middle of a task explains the cost in detail. If you find that a task needs Opus, the lowest-cost moment to switch is after /clear, when the conversation is empty. The page on /clear, /compact and /rewind covers that command.

Effort is easier to change. On Opus 5.5, Sonnet 5.5 and Fable 5.1, with an API key (a pay-per-token account) or a Claude subscription, a change of effort keeps the cache. On most other models, and on Amazon Bedrock or Google Cloud's Agent Platform, it does the same as a model switch [6].

Our position

Set sonnet as your saved default, and leave its effort at medium. Change to opus for a task that is an architecture decision, and make that change at the start of a session. Put search and test runs on Haiku through a subagent. Keep fast mode off. Use Fable for long tasks that you give to Claude as one piece of work.

Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. Every price and every statement about model quality on this page is Anthropic's own, read on 4 October 2026.

Common questions

Which model should I use in Claude Code for everyday coding?

Use Sonnet for everyday coding in Claude Code. Anthropic's documentation, read on 4 October 2026, says Sonnet handles most coding tasks well and costs less than Opus, and its alias table describes `sonnet` as the model for daily coding tasks. On the Anthropic API the alias selects Sonnet 5.5, which is listed at $2 per million input tokens and $10 per million output tokens. Run `/model sonnet` to switch and save it as your default.

What is the default model in Claude Code?

The default model in Claude Code is Opus 5.5 on Pro, Max, Team and Enterprise plans, on the Anthropic API, and on Claude Platform on AWS, Amazon Bedrock and Google Cloud's Agent Platform. On Microsoft Foundry it is Sonnet 4.5. That is the list in Anthropic's documentation, read on 4 October 2026. An organization's administrator can set a different default, and a model you save with `/model` replaces the default for your later sessions.

How much more does Opus cost than Sonnet in Claude Code?

Opus 5.5 costs twice what Sonnet 5.5 costs for input and for output, by our arithmetic on Anthropic's list prices. Anthropic's pricing page, read on 4 October 2026, lists Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens, and Claude Sonnet 5.5 at $2 and $10. A cache read is listed at $0.20 per million tokens on both models, so the difference comes from new input and from output.

What is an effort level in Claude Code?

An effort level in Claude Code is the setting for how much the model reasons before it replies. Anthropic's documentation, read on 4 October 2026, lists five levels on current Opus, Sonnet and Fable models: `low`, `medium`, `high`, `xhigh` and `max`. Lower effort is faster and costs less on simple tasks, because the model's reasoning is billed as output tokens. Haiku models are absent from Anthropic's list of models that support effort.

What is the default effort level on Opus 5.5 and Sonnet 5.5?

The default effort level on Opus 5.5 and Sonnet 5.5 is `medium`. Anthropic's documentation, read on 4 October 2026, says every other model that supports effort starts at `high`, except Opus 4.7, which starts at `xhigh`. Anthropic states that in its own testing Opus 5.5 at `medium` matches or exceeds Opus 5 at `high`, and it advises starting at `medium` when you move from Opus 5 to Opus 5.5.

How do I change the effort level in Claude Code?

Change the effort level in Claude Code with the `/effort` command. Type `/effort low`, or another level name, to set it directly, or type `/effort` alone to open a slider. Anthropic's documentation, read on 4 October 2026, says pressing Enter saves the level for that model in later sessions, and pressing `s` applies it to the current session only, from Claude Code v2.1.257. You can also start a session with the `--effort` flag.

Can I turn off thinking in Claude Code?

You can turn off thinking in Claude Code on some models, and it stays on in Opus 5.5, Sonnet 5.5 and the Fable models. Anthropic's documentation, read on 4 October 2026, says those models always use extended thinking, and the switch in `/config` shows `Thinking can't be turned off` for them. On other models, press `Option+T` on macOS or `Alt+T` on Windows and Linux, or set `MAX_THINKING_TOKENS=0` on the Anthropic API.

What does MAX_THINKING_TOKENS do in Claude Code?

`MAX_THINKING_TOKENS` sets a fixed allowance of thinking tokens in Claude Code, on models that use a fixed budget. Anthropic's documentation, read on 4 October 2026, gives `MAX_THINKING_TOKENS=8000` as an example. Fable models, Sonnet 5 and later, and Opus 4.7 and later always choose their own amount of reasoning and ignore a positive number. On Opus 4.6 and Sonnet 4.6 the fixed budget applies after you also set `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1`.

What is opusplan in Claude Code?

`opusplan` is a model setting in Claude Code that uses Opus while you are in plan mode and Sonnet while Claude writes the code. Plan mode is the setting in which Claude reads and proposes a plan without editing. Anthropic's documentation, read on 4 October 2026, says each change into or out of plan mode is then a model switch that starts a new prompt cache. It suits a task that you plan once at the start.

Does fast mode cost more in Claude Code?

Yes. Fast mode in Claude Code costs more per token than standard Opus. Anthropic's pricing page, read on 4 October 2026, lists fast mode on Opus 5.5 at $8 per million input tokens and $40 per million output tokens, against $4 and $20 at standard speed. Anthropic describes fast mode as up to 2.5 times faster, calls it a research preview, an early release that may change, and says subscribers pay for it from usage credits.

Which model should a Claude Code subagent use?

A Claude Code subagent that does a simple task should use Haiku. Anthropic's documentation, read on 4 October 2026, advises setting `model: haiku` in the subagent's configuration for simple tasks. Without a setting, a subagent runs on the main conversation's model, so a session on Opus runs its subagents on Opus too. From Claude Code v2.1.242, the `/tasks` command names the model on each subagent's row.

When should I use Fable in Claude Code?

Use Fable in Claude Code for long tasks that you give to Claude as one piece of work, such as finding the original cause of a fault. Anthropic's documentation, read on 4 October 2026, describes the Fable models as suited to tasks too large to finish in one working session. Claude Fable 5.1 is listed at $10 per million input tokens and $50 per million output tokens, five times the Sonnet 5.5 price, so keep it for that work.