Anthropic's Claude Haiku 5.5 Beats Every Rival at $24 per 1,000 Tasks

Parallel Web Systems' latest Search Capability Leaderboard puts Anthropic's new Haiku 5.5 at the top for cost-per-task, 39 times cheaper than Opus 5.5.

·
·
·
Anthropic's Claude Haiku 5.5 Beats Every Rival at $24 per 1,000 Tasks
Read4 min
TypeNews
SubtopicSmall Models
  • Claude Haiku 5.5 tops Parallel's Search Efficiency leaderboard at $24.3 per 1,000 tasks.
  • Scores 62.9 with search versus 23.5 without, a +39.4 lift from web tools.
  • Roughly 39 times cheaper than top-ranked Claude Opus 5.5 ($949, score 75.4).
  • Benchmark combines DeepSearchQA, Humanity's Last Exam, and Parallel's WISER business-research suite.
  • Priced at $0.10/M input and $0.50/M output tokens under 100K context.
  • Claude Sonnet 5 posts the field's largest search lift at +43.8 points.

Claude Haiku 5.5 tops search efficiency at $24.30 per 1,000 tasks

Anthropic’s compact Claude Haiku 5.5 ranks first for cost efficiency on Parallel’s leaderboard. It scored 62.9 on web-grounded reasoning tasks at an estimated model cost of $24.30 per 1,000 queries. The top-scoring model, Claude Opus 5.5, reached 75.4 at $949 per 1,000, giving Haiku one thirty-ninth of the cost with a 12.5-point score gap.

Parallel, a search API company, publishes two views of the results. Search Intelligence measures answer quality when a model can use web search. Search Efficiency includes models that clear the median score of 62.8, then ranks them by estimated model cost. Haiku 5.5 clears that threshold by 0.1 point and costs less than every other eligible model.

A 100-question test with fixed tools

Each model answers 100 questions spanning three evaluation suites, both with and without Parallel’s search tools. Google’s DeepSearchQA covers multi-step research, Humanity’s Last Exam uses expert-written questions that require web evidence, and Parallel’s WISER benchmark focuses on business research.

Parallel fixes the available tool budget and disables code execution, keeping the comparison focused on the model and retrieval system. The resulting score reflects performance within that setup, rather than a general measure of model quality across coding, writing, or other workloads.

Lift measures the difference between a model’s search-enabled and search-free scores. Haiku 5.5 rises from 23.5 without search to 62.9 with it, a gain of 39.4 points. That increase indicates that retrieval and lookup orchestration contributed heavily to its benchmark performance.

Twelve points separate $24 from $949

Leading Search Intelligence results
Model Score Model cost per 1,000 tasks
Claude Opus 5.5 75.4 $949
Claude Fable 5.1 72.7 $1,655
Pareto 26.9 72.5 $184
GPT-6 Astra 70.8 $401
GPT-6.1 Sol 70.4 $130
Claude Haiku 5.5 62.9 $24.30

Xiaomi’s MiMo v2.6 Pro is the next-cheapest model above the median, at $76.50 per 1,000 tasks. That rate is more than three times Haiku’s. Claude Sonnet 5 records the largest search lift in the field, gaining 43.8 points when retrieval is enabled.

Low token rates drive the lead

Anthropic’s pricing page lists Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for requests containing up to 100,000 input tokens. Anthropic describes those rates as roughly one-tenth of Haiku 4.5’s per-token pricing.

OpenRouter’s listing gives the model a one-million-token context window. Haiku 5.5 is also the first Haiku release with adjustable reasoning effort, allowing applications to vary how much reasoning the model applies to each request.

Anthropic targets subagents, summarization, classification, routing, and browser automation, all common high-volume components in agent pipelines. In a vendor-published customer example, HubSpot’s internal result averaged 92.8% across three runs of its CRM task suite. That private evaluation provides another workload-specific data point, though it is not directly comparable with Parallel’s benchmark.

The result depends on Parallel’s pipeline

  • Retrieval affects the ranking. Every search-enabled run uses Parallel’s Search Fast mode, so the results may change with Google, Bing, another search API, or an in-house retrieval stack.
  • Some grading is binary. Humanity’s Last Exam and WISER classify answers as correct or incorrect, giving near misses no partial credit.
  • Recovery costs are excluded. The published estimates omit recovery attempts for failed tasks, while production systems often retry failed searches or model calls.
  • Latency remains substantial. Haiku averaged 162 seconds per task. GPT-6 Astra averaged 83.7 seconds at 16.5 times the model cost, while Opus averaged 512 seconds.
  • The sample is limited. A 100-question evaluation can be sensitive to prompt selection, subject mix, and grading decisions, so teams should confirm the ranking on representative traffic.

Route by difficulty, then measure

Agentic search systems can use Haiku 5.5 as the default retrieval tier and escalate selected requests to a larger model. A practical routing policy could include the following steps:

  1. Send routine lookups to Haiku. Classification, extraction, summarization, and well-scoped research requests fit its cost profile.
  2. Escalate difficult cases. Ambiguous questions, conflicting sources, failed citation checks, and multi-stage synthesis can move to a higher-scoring model.
  3. Evaluate the full pipeline. Track task success, citation quality, token use, retry rates, and tail latency with the same retrieval provider used in production.

At the leaderboard’s average rate, 10 million Haiku research tasks would cost about $243,000 in model usage. The same volume would cost approximately $765,000 with MiMo v2.6 Pro or $9.49 million with Opus 5.5. Those estimates exclude search fees, retries, orchestration infrastructure, and differences in production prompt length.

Teams currently using Haiku 4.5 for classification or extraction have a clear candidate for evaluation because Haiku 5.5 combines lower nominal token rates with stronger search performance on Parallel’s test. A controlled production trial should determine whether its task accuracy, latency, and total operating cost hold under the application’s own traffic.

Trending
  • No trending articles

Comments

avatar

Next Reads