How to give a Claude Code subagent a smaller model
To give a Claude Code subagent a smaller model, set the `model` field in its definition file, for example `model: haiku`, which Anthropic's cost page recommends for simple subagent tasks. A subagent is a second copy of Claude that works in its own context window and sends its own billed requests. Without a `model` field, a subagent inherits the main conversation's model, so switching the session to Opus moves those subagents to Opus too. To put every subagent on one model, set `CLAUDE_CODE_SUBAGENT_MODEL` together with `CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1`, which requires Claude Code v2.1.257 or later. On Anthropic's pricing page, read on 4 October 2026, Claude Haiku 4.5 is listed at $1 per million input tokens against $4 for Claude Opus 5.5.
Published October 4, 2026. Editorial.
Key takeaways
- Claude Code picks a subagent's model in a fixed order: the model passed for that invocation, the definition's model field, the CLAUDE_CODE_SUBAGENT_MODEL variable, and then the main conversation's model, by Anthropic's documentation read on 4 October 2026.
- A subagent without a model field inherits the session model, so a /model switch to Opus also moves that subagent's research and test runs to Opus.
- CLAUDE_CODE_SUBAGENT_MODEL is a default that a definition overrides; adding CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1 applies one model to every subagent, teammate and workflow agent, and requires Claude Code v2.1.257 or later.
- The built-in Explore and Plan subagents run on the main conversation's model and skip CLAUDE.md, and a project subagent named Explore with model: haiku replaces the built-in one.
- On Anthropic's pricing page read on 4 October 2026, input is $1 per million tokens on Claude Haiku 4.5, $2 on Claude Sonnet 5.5 and $4 on Claude Opus 5.5, with output at $5, $10 and $20.
Claude Code is Anthropic's coding tool. It runs an AI model that reads files, runs commands and edits code from a request you type in a terminal. A subagent is a second copy of Claude that takes one task into its own context window, which is the full set of text a model reads on each request, and returns a summary. Every request the subagent sends is billed in tokens, the pieces of text the model reads and writes, at the price of whichever model the subagent runs on.
That last clause is the subject of this page. The previous page in this guide, when a subagent saves tokens, showed that a subagent's work is paid in full. This page shows how to make that work cheaper by giving the subagent a smaller model, using the settings in Anthropic's documentation read on 4 October 2026. It is part of the guide to the cost of agents and unattended runs in Claude Code.
The model field in a subagent definition
A custom subagent is a Markdown file with a block of settings at the top, called frontmatter, followed by the subagent's system prompt. Claude Code reads these files from .claude/agents/ in the project and ~/.claude/agents/ in your home folder, and detects an edit within a few seconds, so the next delegation uses the new definition [2].
The model field in that frontmatter sets the subagent's model. Anthropic's documentation lists three kinds of value [2]:
- A model alias:
sonnet,opus,haikuorfable. An alias is a short name that points to the current version of a model family. - A full model ID, such as
claude-opus-5-5orclaude-sonnet-5, which accepts the same values as the--modelflag. inherit, which uses the same model as the main conversation.
Anthropic's cost page gives the recommendation in one line: "For simple subagent tasks, specify model: haiku in your subagent configuration." [1] The subagent page lists the same idea among the reasons to use subagents at all: they "control costs by routing tasks to faster, cheaper models like Haiku" [2].
The order Claude Code uses
A subagent's model can come from four places, and the documentation gives the order in which Claude Code checks them [2]:
- The
modelparameter Claude passes for that one invocation. - The definition's
modelfrontmatter, whereinheritselects the main conversation's model. - The
CLAUDE_CODE_SUBAGENT_MODELenvironment variable, when set to a model alias or model ID. - The main conversation's model.
Before Claude Code v2.1.251, the environment variable came first and overrode the parameter and the frontmatter, including model: inherit [2]. Setting the variable to inherit is the same as leaving it unset; before v2.1.196 that value forced subagents onto the main conversation's model [2].
Two details change what an alias means. When the per-invocation parameter or the frontmatter names a family alias such as opus, and the main conversation already runs a model of that family, the subagent runs on the main conversation's exact model, including any [1m] suffix for the extended context window [2]. An alias in CLAUDE_CODE_SUBAGENT_MODEL always resolves to the version the alias points to, even when it names the main conversation's family [2].
A per-invocation model also applies when the subagent is resumed or sent a follow-up message, so the subagent stays on that model. Before v2.1.211, resuming discarded the value [2].
What a switch to Opus does to inheriting subagents
Step four in that order is the one that costs money without anyone choosing it. A subagent with no model field, started while no other source applies, runs on the main conversation's model. Anthropic's model configuration page describes the result: "If you switch models with /model, the switch also reaches subagents that inherit the main conversation's model, because Claude Code resolves their model from the one your session is using when Claude starts them. Switch to Opus before Claude delegates research or test runs to one of them, and that work runs on Opus too." [3]
The cost page repeats the advice beside its model guidance: "Sonnet handles most coding tasks well and costs less than Opus. Reserve Opus for complex architectural decisions or multi-step reasoning." A switch to Opus "also applies to the subagents that inherit your session's model" [1]. The fix is the same in both places: "To keep a custom subagent on a smaller model, set model in its definition." [3] The page on switching model or effort in the middle of a task covers what the switch does to the main conversation's cache.
Run every subagent on one model
The environment variable is a default. "CLAUDE_CODE_SUBAGENT_MODEL is a default, so a subagent's definition or a model Claude passes still takes precedence over it." [2] To apply one model to every subagent, agent team teammate and workflow agent, Anthropic's documentation says to also set CLAUDE_CODE_SUBAGENT_MODEL_FORCE to 1, which requires Claude Code v2.1.257 or later [2]. The documentation's own example puts both in the env block of a settings file:
{
"env": {
"CLAUDE_CODE_SUBAGENT_MODEL": "haiku",
"CLAUDE_CODE_SUBAGENT_MODEL_FORCE": "1"
}
}
The two variables combine in two ways [2]. With both set, subagents run on the model in CLAUDE_CODE_SUBAGENT_MODEL. With only the force variable set, subagents run on the main conversation's model, except that the built-in Explore subagent runs on the model listed for it in the built-in subagents section. While the force variable is on, Claude Code ignores the model field in subagent definitions and Claude cannot pass a model when it starts a subagent. Two kinds of subagent still run on the main conversation's model: a fork, and a skill (a packaged set of instructions for one task) that runs in a subagent with model: inherit [2].
To confirm the setting took effect, run /tasks while a subagent is running; the subagent's row shows the model it runs on, and the effort level when the definition sets one. That display requires Claude Code v2.1.242 or later [2]. An administrator can place the same env block in managed settings, which the page on settings an administrator can enforce describes.
The built-in Explore and Plan agents
Claude Code includes built-in subagents that Claude uses on its own. Anthropic documents the model of each [2]:
- Explore, a read-only agent for searching a codebase, runs on the main conversation's model. When the main conversation runs Fable, Explore runs on the model the
opusalias resolves to if you connect with a Claude subscription, a Console account or an LLM gateway throughANTHROPIC_BASE_URL, and stays on the main conversation's model on Amazon Bedrock, Google Cloud's Agent Platform, Microsoft Foundry, Claude Platform on AWS or a Claude apps gateway. - Plan, the research agent used in plan mode, inherits from the main conversation unless
CLAUDE_CODE_SUBAGENT_MODELis set and forced onto every subagent. Plan mode is the mode in which Claude reads the code and proposes an approach before it edits anything. - General-purpose uses the
CLAUDE_CODE_SUBAGENT_MODELmodel if set and nothing else assigns one, otherwise the main conversation's model. - Among the helper agents, statusline-setup runs on Sonnet and claude-code-guide runs on Haiku.
Explore and Plan also skip your CLAUDE.md files, the instruction files you write for Claude, and the git status snapshot, "to keep research fast and inexpensive" [2]. Setting CLAUDE_CODE_SUBAGENT_MODEL by itself leaves both unchanged [2]. There is a documented way to put exploration on a smaller model: "A user or project subagent named Explore overrides the built-in and keeps its own model field, so define one with model: haiku to run exploration on a lower-cost model." [2]
Two more facts about what the subagent inherits. As of v2.1.198, a subagent inherits the main conversation's extended thinking configuration, and there is no per-subagent thinking setting [2]. The effort frontmatter field, which sets how much the model reasons before it answers, defaults to the session level and can be set to low, medium, high, xhigh or max, with the available levels depending on the model; a maxEffortLevel setting or an organisation effort cap still limits it [2][3]. The page on which model and effort level to use covers those levels for the main conversation.
Suggested model by task
The table pairs the kinds of task Anthropic's documentation names with the model its documentation suggests. The final column says where the suggestion comes from.
| Subagent task | Suggested model | Reason, by Anthropic's documentation |
|---|---|---|
| Simple, well-defined work: run the tests, report the failures | haiku |
The cost page says to specify model: haiku for simple subagent tasks [1] |
| Searching and reading a codebase | A project subagent named Explore with model: haiku |
The documented way to run exploration on a lower-cost model [2] |
| Coordination between agents, as in an agent team | sonnet |
The cost page says Sonnet "balances capability and cost for coordination tasks" [1] |
| Most coding tasks | sonnet |
"Sonnet handles most coding tasks well and costs less than Opus" [1] |
| Complex architectural decisions or multi-step reasoning | opus |
The cost page says to reserve Opus for these [1] |
| The hardest and longest-running tasks | fable |
The alias table describes the Fable model this way [3] |
| A task that needs the parent's exact context | inherit, or a fork |
A fork runs on the main session's model and reads its cache [2] |
The last row sets a limit on the whole idea. A fork inherits the main session's model as part of inheriting its prompt cache, so the saving from a smaller model and the saving from a shared cache are available on different subagents, never on the same one [2].
What the smaller model costs, in list prices
Anthropic's pricing page, read on 4 October 2026, lists these prices in US dollars per million tokens [4]:
| Model | Input | Cache read | Output |
|---|---|---|---|
| Claude Haiku 4.5 | $1 | $0.10 | $5 |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 |
| Claude Opus 5.5 | $4 | $0.20 | $20 |
| Claude Fable 5.1 | $10 | $0.25 | $50 |
On the Anthropic API, the opus alias resolves to Opus 5.5 and sonnet to Sonnet 5.5, and fable resolves to Fable 5.1 unless you set ANTHROPIC_DEFAULT_FABLE_MODEL, except in Claude apps gateway sessions, where it resolves to Fable 5 [3]. The model configuration page describes the haiku alias as "the fast and efficient Haiku model for simple tasks" and names no version; Claude Haiku 4.5 is the Haiku model on the pricing page that is listed without a retirement note [3][4].
An invented example, with our own arithmetic on those list prices and no cache reads: a subagent that reads 100,000 input tokens and writes 5,000 output tokens costs $0.125 on Haiku 4.5 ($0.10 input plus $0.025 output), $0.25 on Sonnet 5.5 ($0.20 plus $0.05) and $0.50 on Opus 5.5 ($0.40 plus $0.10). At these prices the Opus 5.5 run costs four times the Haiku 4.5 run. Two cautions apply. First, these are API list prices; no documentation page says how tokens count against a subscription allowance. Second, the pricing page says Claude 4.7 and later models use a newer tokenizer that Anthropic puts at 30 percent more tokens for the same text, while Sonnet 4.6 and earlier use the previous one, so the same files can be more tokens on a newer model than on Haiku 4.5 [4].
Our position
Give every subagent a model field, and make the default small. Reveneau recommends model: haiku on subagents that run tests, search logs or fetch documentation, model: sonnet on subagents that write code, and opus only on a subagent whose task is a design decision or multi-step reasoning, which is where the cost page says to reserve Opus. A definition without the field follows the session, and the session follows whoever last ran /model. For a team that wants one answer for every machine, the two environment variables with the force flag give it, with the Explore exception and the fork exception noted above.
The next size above a subagent is an agent team, where each teammate is a full session. The page on what agent teams cost covers that, and what a workflow costs covers runs of many agents from a script.
Reveneau is an AI software development consultancy. All of its code is written by AI and every change must pass an eval suite, a set of automated tests written from the specification, before release, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. Every setting, version number and price on this page is Anthropic's own statement about its own product, read on 4 October 2026.
Common questions
How do I set the model for one subagent in Claude Code?
Set the `model` field in the subagent's definition file, the Markdown file with a settings block (frontmatter) at the top, kept under `.claude/agents/` or `~/.claude/agents/`. Anthropic's documentation, read on 4 October 2026, says the field accepts an alias such as `sonnet`, `opus`, `haiku` or `fable`, a full model ID such as `claude-opus-5-5`, or `inherit`, which selects the main conversation's model. Claude Code detects an edit to the file within a few seconds, and the next delegation uses the new definition.
In what order does Claude Code pick a subagent's model?
Claude Code picks a subagent's model from the first of four sources that applies, by Anthropic's documentation read on 4 October 2026: the `model` parameter Claude passes for that one invocation, the definition's `model` frontmatter, the `CLAUDE_CODE_SUBAGENT_MODEL` environment variable when set to an alias or model ID, and then the main conversation's model. Before Claude Code v2.1.251, the environment variable came first and overrode both the parameter and the frontmatter.
What happens to subagents when I switch the session to Opus with /model?
A `/model` switch to Opus also reaches every subagent that inherits the main conversation's model, because Claude Code resolves their model from the session's model at the moment Claude starts them. Anthropic's documentation, read on 4 October 2026, gives the example: switch to Opus before Claude delegates research or test runs, and that work runs on Opus too. To keep a custom subagent on a smaller model, set `model` in its definition.
How do I run every subagent on Haiku?
Set two environment variables in the `env` block of a settings file: `CLAUDE_CODE_SUBAGENT_MODEL` to `haiku` and `CLAUDE_CODE_SUBAGENT_MODEL_FORCE` to `1`. Anthropic's documentation, read on 4 October 2026, says the first variable alone is a default that a definition's `model` field or a model Claude passes still overrides, and that the second makes it apply to every subagent, agent team teammate and workflow agent. The force variable requires Claude Code v2.1.257 or later.
What does CLAUDE_CODE_SUBAGENT_MODEL_FORCE do?
`CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1` makes Claude Code ignore the `model` field in subagent definitions and stops Claude passing a model when it starts a subagent, by Anthropic's documentation read on 4 October 2026. With `CLAUDE_CODE_SUBAGENT_MODEL` also set, every subagent runs on that model. With only the force variable set, subagents run on the main conversation's model, except the built-in Explore subagent. A fork, and a `context: fork` skill with `model: inherit`, still run on the main conversation's model.
Which model do the built-in Explore and Plan subagents use?
Explore uses the main conversation's model, and Plan inherits from the main conversation unless `CLAUDE_CODE_SUBAGENT_MODEL` is set and forced onto every subagent, by Anthropic's documentation read on 4 October 2026. When the main conversation runs Fable with a Claude subscription, a Console account or an LLM gateway through `ANTHROPIC_BASE_URL`, Explore runs on the model the `opus` alias resolves to. Both skip CLAUDE.md and the git status snapshot. Setting `CLAUDE_CODE_SUBAGENT_MODEL` by itself leaves both unchanged.
How do I check which model a subagent is running on?
Run `/tasks` while the subagent is running. Anthropic's documentation, read on 4 October 2026, says Claude Code names the model on the subagent's row, and adds the effort level when the subagent's definition, or the skill it forked from, sets `effort`. This requires Claude Code v2.1.242 or later. In interactive sessions Claude Code also shows a warning naming the requested model and the model the subagent runs on whenever it substitutes one for the other.
What does Claude Haiku 4.5 cost per million tokens on the Anthropic API?
On Anthropic's pricing page, read on 4 October 2026, Claude Haiku 4.5 is listed at $1 per million input tokens, $1.25 for a five-minute cache write, $2 for a one-hour cache write, $0.10 for a cache read and $5 per million output tokens. Those are API list prices in US dollars. No Anthropic documentation page says how tokens count against a Pro, Max, Team or Enterprise subscription allowance, so the figures apply to pay-per-token billing.
Does a subagent inherit the session's effort level?
Yes by default. The effort level is the setting for how much the model reasons before it answers, and Anthropic's documentation, read on 4 October 2026, says a subagent's `effort` frontmatter field defaults to inheriting the session level. Setting the field to `low`, `medium`, `high`, `xhigh` or `max` overrides the session level while that subagent is active; the available levels depend on the model, and a `maxEffortLevel` setting or an organisation effort cap still limits the level the subagent runs at.
What happens when availableModels blocks the model a subagent asks for?
When the `availableModels` allowlist blocks a family alias such as `opus`, Claude Code runs the subagent on the newest version of that family the allowlist permits, by Anthropic's documentation read on 4 October 2026. For any other blocked value, on providers where that substitution does not operate, or when no version of the family is permitted, the subagent runs on the inherited model instead; if `CLAUDE_CODE_SUBAGENT_MODEL` is set, Claude Code tries that model first under the same rules.
Does a subagent get the same context window size as the main conversation?
A subagent's context window is sized by its own model, by Anthropic's documentation read on 4 October 2026, so delegating to a model with a smaller window gives that subagent the smaller window. One exception: when a family alias such as `opus` in the per-invocation parameter or the frontmatter names the family the main conversation already runs, the subagent runs on the main conversation's exact model, including any `[1m]` suffix, and gets the same extended context window.
Does a subagent use extended thinking?
As of Claude Code v2.1.198, a subagent inherits the main conversation's extended thinking configuration: if thinking is on in your session, it is on for the subagent, and if it is off, it stays off. Anthropic's documentation, read on 4 October 2026, says there is no per-subagent thinking setting. Before v2.1.198, subagents ran with extended thinking disabled regardless of the session. Thinking tokens are billed as output tokens, the most expensive type on every model.