Cursor Adds Claude Haiku 5.5 at 10x Cheaper Rates for Coding Tasks

Anthropic's new small model lands in Cursor at roughly a tenth of Haiku 4.5's price on short requests, with Sonnet 5.5 cache reads also cut in half.

·
·
·
Cursor Adds Claude Haiku 5.5 at 10x Cheaper Rates for Coding Tasks
Read4 min
TypeNews
  • Cursor enabled Claude Haiku 5.5, toggleable under Settings > Models in the editor.
  • Priced at $0.10/M input and $0.50/M output tokens, roughly 10x cheaper than Haiku 4.5 on short requests.
  • Long contexts above 100k input tokens jump to $0.50/M input and $2.50/M output.
  • Sonnet 5.5 cache reads also dropped from $0.20/M to $0.10/M, cutting agent-loop bills.
  • Haiku 4.5 scored over 73% on SWE-bench Verified, setting a high bar for the new release.
  • Best fit for sub-agent execution, inline edits, classification, and high-frequency chat turns.

Cursor adds Claude Haiku 5.5 at one-tenth the short-context price

Cursor has added Claude Haiku 5.5 to its editor, giving developers access to Anthropic’s new small-tier model before Anthropic has published a full specification. Cursor lists short-context prices at one-tenth of Haiku 4.5’s rates, making the model a cheaper option for frequent, narrowly scoped coding tasks.

Developers can enable the model under Settings > Models and select it from Cursor’s model picker. If it does not appear, update Cursor before checking the settings again.

Short prompts get the largest discount

Cursor lists two pricing tiers based on input length. Input tokens include instructions, conversation history, retrieved files, and other context sent with a request.

Input length Input price Output price Difference from Haiku 4.5
Up to 100,000 tokens $0.10 per million tokens $0.50 per million tokens 10 times cheaper
More than 100,000 tokens $0.50 per million tokens $2.50 per million tokens 2 times cheaper
Haiku 4.5 $1 per million tokens $5 per million tokens Baseline

A request containing 50,000 input tokens and producing 10,000 output tokens would cost about $0.01 at the short-context rate. The same token counts cost about $0.10 with Haiku 4.5. Requests exceeding 100,000 input tokens receive a smaller discount, which limits the savings for large repositories and long agent histories.

Cursor also reduced Sonnet 5.5 cache-read pricing from $0.20 to $0.10 per million tokens. Cached reads let an agent reuse previously processed prompt prefixes or files, so the reduction can lower costs during long sessions that repeatedly reference the same codebase context.

Why a cheaper Haiku fits an IDE

Haiku models target low-latency, high-volume work such as focused edits, file searches, diff summaries, request classification, and bounded sub-agent tasks. An IDE can generate many of these calls during a single coding session, making per-token cost more consequential than it is for occasional chat requests.

Agent systems can also divide work by model capability. A Sonnet or Opus model can plan a repository-wide change while several Haiku workers inspect files, search for symbols, run targeted analyses, or prepare small patches. Lower worker costs reduce the total price of that fan-out pattern while preserving the larger model for tasks that require broader reasoning.

The specification still has gaps

Anthropic has yet to publish a complete Haiku 5.5 model card, leaving its context window, maximum output, knowledge cutoff, safety classification, and benchmark results unconfirmed. Cursor’s availability and pricing therefore provide more information about deployment economics than underlying capability.

Haiku 4.5 provides a useful baseline, according to Anthropic’s Haiku 4.5 announcement:

  • A 200,000-token context window and a 64,000-token maximum output
  • A February 2025 knowledge cutoff
  • Support for extended thinking and Computer Use
  • An AI Safety Level 2 classification
  • A vendor-reported score above 73% on SWE-bench Verified, a benchmark built from real GitHub issues

Those Haiku 4.5 specifications should not be assumed for Haiku 5.5 until Anthropic publishes the new model’s documentation. Cursor provides separate CursorBench results for comparisons based on editor tasks.

Where the lower rate pays off

Workload Why Haiku 5.5 may fit When to use a larger model
Inline edits and small refactors Requests are narrow, frequent, and latency-sensitive. The change spans many files or requires architectural judgment.
Agent sub-tasks Searches, file reads, and small patches can run in parallel. The task requires coordination across dependent changes.
Classification and routing Short outputs and high request volume favor the lower rate. Categories are ambiguous or depend on extensive context.
Documentation and support queries Retrieved context can keep each answer focused. Answers require deep synthesis across a large corpus.
Repository-wide planning Haiku can gather facts for a separate planner. Sonnet or Opus should own the plan and resolve trade-offs.

Teams evaluating a default-model change can compare latency, accepted edits, retry rates, token use, and total task cost on a representative set of repository tasks. The higher long-context tier makes prompt size especially important, while retries can erase savings from a cheaper initial call.

Cheaper workers reshape agent routing

Haiku 5.5’s short-context pricing favors systems that reserve expensive models for planning and route bounded execution steps to smaller workers. That architecture becomes less economical as prompts cross 100,000 input tokens or weak outputs require repeated attempts, so effective routing depends on task size, context length, and measured success rates.

Trending
  • No trending articles

Comments

avatar

Next Reads