Anthropic's Claude Haiku 5.5 Jumps 89 Spots to Rank Third in Coding

Vals AI benchmarks reveal Anthropic's new small model climbs to #3 on Vibe Code Bench by burning far more reasoning tokens per task.

·
·
·
Anthropic's Claude Haiku 5.5 Jumps 89 Spots to Rank Third in Coding
Read4 min
TypeNews
  • Claude Haiku 5.5 scores 90.44% on Vibe Code Bench v1.1, ranking 3rd of 110 models.
  • Base pricing is $0.10 input / $0.50 output per million tokens, 10x cheaper than Haiku 4.5.
  • Prices jump 5x once a request exceeds 100k tokens of context.
  • On Legal Research it averages 59 steps per task versus 17 for Haiku 4.5, writing 15x more tokens.
  • Ranks #16 on the Vals Index at 54.31% and $2.99 per test.
  • Ships with a 1M context window and 128k max output tokens.

Claude Haiku 5.5 climbs coding benchmarks by spending more reasoning tokens

Vals AI evaluation data shows Claude Haiku 5.5, Anthropic’s lower-cost Claude tier, outperforming Haiku 4.5 on every benchmark that included both models. It ranks third among 110 models on Vibe Code Bench v1.1 while costing about one-fifth as much as Sonnet 5.5 in that test. The gains come with heavy extended-thinking output, which can raise per-task costs and narrow the model’s advantage on long-context workloads.

Code lifts Haiku into the top three

Vibe Code Bench v1.1 asks models to build complete applications from plain-English specifications. Haiku 5.5 scored 90.44%, placing third among 110 tested models and climbing 89 positions from Haiku 4.5.

Vals reports that Haiku 5.5 passed every test for 21 of the benchmark’s 50 applications, compared with 24 for Sonnet 5.5. Haiku completed those runs at about one-fifth of Sonnet’s cost and in roughly half the time.

Across the Vals Index, a cross-benchmark ranking, Haiku 5.5 placed 16th among 45 models with a score of 54.31%. Its average cost of $2.99 per test made it the second-cheapest model in the top 16.

Vals AI results for Claude Haiku 5.5
Benchmark Score Rank
Vibe Code Bench v1.1 90.44% 3 of 110
ProofBench v1.1 86.00% 10 of 48
Tax Agent Bench 62.20% 30 of 67
EMB 59.02% 34 of 72
Finance Agent v2 54.14% 24 of 76
IOI 47.28% 27 of 42
Legal Research Bench 43.27% 20 of 75
Terminal-Bench 4.0 35.35% 12 of 45
Harvey’s Legal Agent Benchmark 1.25% 54 of 76

Vals’ scores reflect its prompts, tool harnesses, and run settings. Production results can vary with task design, latency limits, tool reliability, retry policies, and acceptable failure rates.

Longer reasoning traces drive the gain

On Legal Research Bench, Vals recorded an average of 59 agent steps per task for Haiku 5.5, compared with 17 for Haiku 4.5. An agent step represents a model turn or tool interaction. The newer model generated roughly 15 times as much output, and extended-thinking tokens accounted for about 80% of that output.

The same token-heavy pattern appeared against Sonnet 5.5 and Opus 5.5. On the legal research, tax, and finance benchmarks, Haiku generated more output than both larger Claude tiers, mostly through extended thinking. It used fewer agent steps than Sonnet 5.5 overall, with longer deliberation inside each model turn. Those thinking tokens contribute to metered output usage.

The 100,000-token threshold changes the economics

Haiku 5.5’s listed base rates are $0.10 per million input tokens and $0.50 per million output tokens, one-tenth of Haiku 4.5’s respective $1 and $5 rates. Requests exceeding 100,000 tokens enter a pricing tier that raises both rates fivefold.

Claude Haiku 5.5 token pricing
Context tier Input per million tokens Output per million tokens
Up to 100,000 tokens $0.10 $0.50
Over 100,000 tokens $0.50 $2.50

Requests in the higher tier also cost $0.05 per million tokens for cache reads and $0.625 per million tokens for cache writes. Most of Vals’ Legal Research requests crossed the 100,000-token threshold.

The higher tier, combined with Haiku 5.5’s larger reasoning output, pushed its per-test cost above Haiku 4.5 on every shared benchmark. Developers evaluating agent workloads therefore need request-level estimates that include context length, thinking tokens, cache operations, retries, and tool loops.

Inside Vals’ test harness

  • Context window: 1 million tokens
  • Maximum output: 128,000 tokens
  • Temperature: 1.0
  • Sampling: Default top-p and top-k
  • Compute effort: Maximum
  • Fallback rate: 0.00%
  • Refusal rate: 0.22%

Anthropic’s Claude model guide provides the broader model specifications, while the Vals results describe the configuration used for these benchmark runs.

Bounded workloads preserve the price advantage

Workload selection signals from the benchmark data
Workload Selection signal Primary constraint
Specification-to-app generation Strong benchmark candidate Validate framework support, correctness, and repair rates
Short coding and utility tasks Low base token rates Keep prompts and output lengths bounded
Long legal or financial agents Higher per-task cost risk Large reasoning traces and the higher context tier
Legal agents resembling Harvey’s benchmark Weak benchmark result 1.25% score and a rank of 54 among 76 models

Cost modeling should start with representative production traces that include prompt length, answer tokens, extended-thinking tokens, cache operations, tool calls, retries, and the share of requests crossing 100,000 tokens. Vals’ data supports testing Haiku 5.5 for bounded coding workloads. Long-running agents require per-task estimates because heavy reasoning output and the higher pricing tier can consume the savings from its low base rate.

Trending
  • No trending articles

Comments

avatar

Next Reads