AI NewsModels & agentsAnnouncement
Anthropic launches Claude Haiku 5.5 at $0.10 per million input tokens
Anthropic released Claude Haiku 5.5, dropping the input price to $0.10 per million tokens for requests up to 100,000 tokens, which the company says is 90 percent below Haiku 4.5.

Image: Anthropic
Why it mattersA high-volume task that was borderline at Haiku 4.5 pricing is now ten times cheaper per input token, so classification, summarisation and subagent calls teams had moved to batch runs to afford are back in reach for real-time use.
A high-volume classification pass that cost $100 a day on Haiku 4.5 now costs about $10 at the same traffic. Anthropic released Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for requests up to 100,000 tokens, and the company says this is 90 percent lower than Haiku 4.5 at that request size, and 50 percent lower for requests over 100,000 tokens.
Above 100,000 input tokens the price steps up to $0.50 input and $2.50 output per million. Cache reads come in at $0.01 and $0.05 per million for the two tiers, and cache writes at $0.125 and $0.625. The model is available now on the Anthropic API, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Azure, per Anthropic.
What Anthropic says it does better
Anthropic names four benchmark numbers in the launch note, each as its own figure against Haiku 4.5. On GDPval-AA v2.1, the company reports 1620 for Haiku 5.5 against 735 for Haiku 4.5. On the OSWorld 2.1 offline subset, 72.4 percent against 15.7 percent. On Humanity's Last Exam with no tools, 45.9 percent against 10.2 percent. On Terminal-Bench 4.0, 39.2 percent against 0.0 percent.
Those are the company's own figures and have not been independently reproduced yet, so read the Terminal-Bench and OSWorld scores against what the Haiku line was asked to do before: Haiku has not been the model teams used for agentic coding or desktop control. The move from 0.0 to 39.2 on Terminal-Bench is the kind of change that makes it worth re-evaluating which tasks route to Haiku against the mid-tier model.
The use cases Anthropic names
The launch note lists the work Anthropic is pointing Haiku 5.5 at: "summaries, compactions, database queries, and classification requests", plus "live customer support and browser use" and "subagent work". That last use is the one that matters most for a coding or research agent, where the top-tier model plans and a cheaper model does the many small calls the plan expands into.
Price caps for a subagent pattern depend on Haiku's price more than on any other model in the mix. A plan that spawns 80 Haiku calls per task was already affordable at Haiku 4.5; at Haiku 5.5 the input cost on a 1,000-token-per-call pattern is one tenth of a cent for the batch, and the budget question shifts to the output tokens the subagents produce.
What stays the same
The 100,000-token price break is the same cutoff Anthropic has used on the rest of the line, so a long-context job still pays the long-context rate. The model is not free on any tier and the launch note does not state a context window. Workflows that currently depend on cache reads to hold a system prompt or a repository index in place benefit most from the $0.01 cache tier, because at that price a cache miss on a long prefix starts to look like the dominant cost on a short request.
For a team deciding whether to migrate production traffic, the useful number is the ratio of Haiku 5.5 output to Haiku 4.5 output on your own eval, measured on your real prompts. A classifier whose output stayed at parity is a straight tenfold price cut on inputs. A classifier whose output moved by three percentage points, in either direction, is a different decision.
Source
Claude Haiku 5.5, Anthropic. Press coverage: Anthropic launches Haiku 5.5 at a much lower price, The New Stack.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.