Anthropic's Claude Haiku 5.5 Slashes Costs 75% and Adds Adjustable Reasoning

Anthropic's smallest model gets a 1M context window, adjustable effort, and roughly 75% lower cost per token than Haiku 4.5.

·
·
·
Read5 min
TypeNews
TopicLlms · Api
  • Anthropic released Claude Haiku 5.5, roughly 75% cheaper on average than Haiku 4.5.
  • Pricing starts from $0.10/MTok input and $0.50/MTok output, scaling with effort.
  • First Haiku with an adjustable effort setting to trade cost against intelligence per call.
  • Context window jumps to 1M tokens with 128K max output; model ID claude-haiku-5-5.
  • Available on Anthropic API, AWS, Google Cloud, and Azure.
  • Sonnet 5.5 cache reads halved to $0.10/MTok, cutting long-running workloads about 20%.

Claude Haiku 5.5 cuts costs and adds adjustable effort

Anthropic has released Claude Haiku 5.5, the latest version of its small, fast model tier. It adds a 1 million-token context window, supports up to 128,000 output tokens, and lets developers adjust reasoning effort for each request. Anthropic says average workloads cost about 75% less than Haiku 4.5. The model is available through the Anthropic API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Azure.

A lower floor with variable pricing

Anthropic lists Haiku 5.5 at rates starting at $0.10 per million input tokens and $0.50 per million output tokens in its model documentation. Those rates apply at the lowest effort setting; higher reasoning effort increases token use and cost.

Model Input tokens Output tokens
Claude Haiku 5.5 From $0.10 per million From $0.50 per million
Claude Haiku 4.5 $1 per million $5 per million

At the lowest effort setting, both input and output rates are 90% below Haiku 4.5. Anthropic estimates a smaller average reduction of about 75% across typical workloads because requests that use more reasoning cost more. Teams evaluating the model should compare total request cost at the effort levels their applications require.

A much larger working envelope

The model expands Haiku’s context window from 200,000 to 1 million tokens while raising maximum output to 128,000 tokens. Its default reasoning effort is medium.

Model ID claude-haiku-5-5
Context window 1 million tokens
Maximum output 128,000 tokens
Reasoning mode Adaptive effort, defaulting to medium
Reliable knowledge cutoff June 2026

One model, several compute budgets

Haiku 5.5 is the first Haiku release with adjustable effort, allowing applications to assign a different reasoning budget to each request. A classification call can use the lowest setting, while a coding subagent can use high effort when a task requires more analysis.

A per-request effort setting can simplify routing for workloads that previously moved borderline tasks between Haiku and Sonnet. Applications can keep a single model integration and tune the quality, latency, and cost trade-off through one parameter. Production tests should measure accuracy, response time, token use, and tool-call reliability at each relevant setting.

Tuned for frequent, short requests

Anthropic positions Haiku 5.5 for high-volume workloads such as summarization, classification, extraction, request routing, customer support, and browser or computer control. These applications often issue many sequential calls, so modest reductions in per-call latency and cost can accumulate across a session.

Anthropic says requests below 100,000 tokens represented about 90% of traffic to Haiku 4.5. Haiku 5.5’s pricing targets that common range, while the larger context window supports occasional tasks involving repositories, long transcripts, document collections, or extended agent histories.

In multi-agent coding systems, Opus 5.5 or Sonnet 5.5 can serve as the planner while parallel Haiku 5.5 subagents read files, search code, run tests, and inspect results. The 1 million-token window also allows those subagents to receive a larger share of the planner’s working context.

Capability claims still need benchmarks

Anthropic reports gains over Haiku 4.5 in coding, computer use, and knowledge work, although its launch thread does not include a full benchmark table. Developers will need workload-specific evaluations to determine how those gains change with effort level and whether Haiku 5.5 can replace Sonnet for particular tasks.

The company also reports fewer instances of misaligned behavior across nearly all of its alignment evaluations. That claim is especially relevant to agents that execute tool calls or interact with interfaces repeatedly, where small errors can compound over many steps.

Haiku 5.5 follows Opus 5.5 and Sonnet 5.5 in Anthropic’s current model family. Anthropic previously reported that Sonnet 5.5 generated output more than 30% faster, cost up to 30% less per task, and approached Opus 5.5 on knowledge-work benchmarks. Its reported Terminal-Bench result, which measures performance on terminal-based tasks, rose from 10.3% to 70.6%.

Sonnet cache reads also get cheaper

Anthropic has also halved Claude Sonnet 5.5 cache-read pricing to $0.10 per million tokens. The company estimates that the change reduces costs by about 20% for most long-running workloads.

Prompt caching lets applications reuse shared context, such as a system prompt, repository snapshot, or long transcript, without processing the same tokens at full input rates on every call. Applications already using Sonnet 5.5 caching can receive the lower read rate when their provider adopts the updated pricing; actual savings depend on cache-hit frequency and the amount of reused context.

Best candidates for migration

  1. Current Haiku 4.5 users: Test Haiku 5.5 as a direct model replacement, beginning with the default effort setting and comparing cost, latency, and output quality.
  2. Teams using Sonnet for structured tasks: Re-evaluate classification, extraction, and routing workloads with Haiku 5.5 at medium and high effort.
  3. Agent developers: Consider Haiku 5.5 for parallel file inspection, search, testing, and browser actions dispatched by a larger planner model.
  4. Long-context applications: Test whether the 1 million-token window reduces chunking or retrieval overhead while remaining within latency and cost targets.

Migration checklist

  • Update the model ID to claude-haiku-5-5 and confirm that the selected provider or SDK supports it.
  • Set effort explicitly for predictable behavior, since the default is medium.
  • Benchmark representative requests at each planned effort level.
  • Track input, output, cached, and reasoning-related token costs separately.
  • Validate tool calls, structured outputs, and agent stopping behavior before expanding traffic.
  • Roll out gradually and retain a fallback model for tasks that miss quality or latency thresholds.

Haiku 5.5 gives production systems a wider range of cost and reasoning options under one model ID. Its practical value will depend on whether adjustable effort lets each workload meet its quality target while reducing total cost and routing complexity.

Trending
  • No trending articles

Comments

avatar

Next Reads