Anthropic launches Claude Haiku 5.5, around 75% cheaper to run on average

Anthropic launches Claude Haiku 5.5, around 75% cheaper to run on average

Anthropic has released Claude Haiku 5.5, which it calls the cheapest, fastest and most capable small model it has ever released. It is available now on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure. On the Claude Platform, developers can start with the model name claude-haiku-5-5, and a migration guide is linked.

Anthropic positions Haiku 5.5 for high-volume, cost-sensitive work: quick and repetitive jobs such as summaries, compactions, database queries and classification requests. It is meant to pair with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. Because it is also Anthropic's fastest model to date, the company says it suits speed-sensitive uses like live customer support and browser use. A footnote qualifies the speed claim: Haiku 5.5 is the fastest model at each model's standard speed, but it runs less quickly than the Opus models in Fast Mode.

On price, Anthropic says Haiku 5.5 costs around 75% less to run than Haiku 4.5, on average. The footnote explains the arithmetic. The model is priced 90% lower than Haiku 4.5 for requests up to 100,000 tokens and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the shorter category. The calculation also accounts for a new tokenizer, similar to those of Sonnet 5.5 and Opus 5.5, which makes Haiku 5.5 use slightly more tokens per task.

Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, so users can choose whether to optimise for cost or for intelligence, as with Anthropic's other models. The announcement shows benchmark charts, including results on three benchmarks at each effort setting, and says early customers reported results consistent with the performance and cost improvements shown.

On safety, Anthropic says Haiku 5.5 shows major improvements across almost all of its alignment evaluations relative to Haiku 4.5, with far fewer instances of misaligned behavior and a lower willingness to cooperate with misuse. Its cybersecurity safeguards are more restrictive than Haiku 4.5's but somewhat less restrictive than those on other recent models. They permit a wider range of defensive tasks than the Sonnet 5.5 safeguards, but still block penetration testing and other techniques more likely to be used by attackers. The biology safeguards match those of Sonnet 5, Sonnet 5.5 and Opus 5: research biology questions are allowed, while requests judged likely to cause harm are restricted. Organisations doing wider-ranging biology and cyber work can apply to the Life Sciences Verification Program and the Cyber Verification Program.

Two further changes came with the launch. First, starting the day of the announcement, Sonnet 5.5 cache reads cost 50% less: $0.10 per million tokens rather than $0.20. Since cache reads make up a large share of token consumption, Anthropic says this cuts the cost of Sonnet 5.5 on most agentic tasks by around 20%. Second, this week Anthropic will roll out a monthly Claude Platform API credit to all Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 for Team, pooled across users. The credits can be used on any Anthropic model and are meant for experimenting with tools, apps and agents that call the API.

Anthropic also stresses a limit. Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding of the kind measured by Terminal-Bench 4.0. Haiku 5.5 is best for more narrowly scoped tasks, such as compaction, summarisation or subagent work, that might previously have been too costly. Finally, the Claude Python and TypeScript SDKs are being updated with beta support for computer use and browser use, tasks Anthropic says suit Haiku 5.5 given its speed, capability and price.

Key facts

  • Claude Haiku 5.5 costs around 75% less to run than Haiku 4.5 on average: 90% lower for requests up to 100,000 tokens and 50% lower above that, adjusted for a tokenizer that uses slightly more tokens per task.
  • It is the first Haiku-class model with an adjustable effort setting, and is available now on AWS, Google Cloud and Microsoft Azure as claude-haiku-5-5.
  • Sonnet 5.5 cache reads drop from $0.20 to $0.10 per million tokens, which Anthropic says cuts Sonnet 5.5 cost on most agentic tasks by around 20%.
  • Max 5x subscribers get $100 per month in Claude Platform API credit, Max 20x get $200, and Team up to $500 pooled, usable on any model.
  • Anthropic says Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding such as Terminal-Bench 4.0; Haiku 5.5 targets narrowly scoped, high-volume tasks.

Why it matters

The small-model tier is where high-volume workloads live, and Anthropic is cutting its cost sharply: around 75% on average versus Haiku 4.5. The company argues this brings within reach tasks that were previously cost-prohibitive, such as compaction, summarisation and subagent work. The first Haiku-class adjustable effort setting also gives small-model users the cost-versus-intelligence dial that the larger models already had. The Sonnet 5.5 cache read cut, around 20% on most agentic tasks, extends the price move beyond the new model.

Who it affects

Developers running high-volume, cost-sensitive pipelines such as classification, database queries, summaries and live customer support are the main audience. Teams building coding agents may use Haiku 5.5 as a subagent alongside Opus 5.5 or Sonnet 5.5. Sonnet 5.5 users with cache-heavy agentic workloads get a lower bill from the cache read cut. Max 5x, Max 20x and Team subscribers receive the new monthly API credits, and organisations needing broader biology or cyber capabilities can apply to Anthropic's verification programs.

How to use it

Use the model name claude-haiku-5-5 on the Claude Platform; it is also available on Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic points to a migration guide. The effort setting lets you trade cost against intelligence per request. The Python and TypeScript SDKs are gaining beta support for computer use and browser use. Max and Team subscribers should see their monthly credits arrive this week ($100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team), usable on any Anthropic model. Anthropic recommends Haiku 5.5 for narrowly scoped work and the larger models for complex agentic coding.

How solid is it

This is Anthropic's own first-party announcement, so the performance, speed and safety claims are the company's, not independent measurements. The price arithmetic is laid out in a footnote: 90% lower under 100,000 tokens, 50% lower above, weighted by the share of Haiku 4.5 requests (90%) in the shorter bracket and adjusted for the new tokenizer. The speed claim is qualified: fastest at standard speed, slower than Opus models in Fast Mode. Customer feedback is described as early testing and consistent with the shown improvements. The system card holds the detailed alignment evaluations.

Risks and caveats

The 75% figure is an average that depends on the mix of request lengths and the tokenizer change; it is not a flat price cut, and workloads dominated by requests over 100,000 tokens would see a smaller saving (50% lower per-token price before the tokenizer effect). The roughly 20% Sonnet 5.5 saving applies to most agentic tasks, not all workloads. Haiku 5.5 uses slightly more tokens per task than before. Its cybersecurity safeguards are stricter than Haiku 4.5's and still block penetration testing, which may constrain some security teams. Anthropic itself says Haiku 5.5 is not the right pick for complex agentic coding.

“the cheapest, fastest, and most capable small model we’ve ever released”

— Anthropic, announcement of Claude Haiku 5.5