Anthropic's Claude Haiku 5.5 Jumps 89 Spots to Rank Third in Coding
Vals AI benchmarks reveal Anthropic's new small model climbs to #3 on Vibe Code Bench by burning far more reasoning tokens per task.
- Claude Haiku 5.5 scores 90.44% on Vibe Code Bench v1.1, ranking 3rd of 110 models.
- Base pricing is $0.10 input / $0.50 output per million tokens, 10x cheaper than Haiku 4.5.
- Prices jump 5x once a request exceeds 100k tokens of context.
- On Legal Research it averages 59 steps per task versus 17 for Haiku 4.5, writing 15x more tokens.
- Ranks #16 on the Vals Index at 54.31% and $2.99 per test.
- Ships with a 1M context window and 128k max output tokens.
Claude Haiku 5.5 climbs coding benchmarks by spending more reasoning tokens
Vals AI evaluation data shows Claude Haiku 5.5, Anthropic’s lower-cost Claude tier, outperforming Haiku 4.5 on every benchmark that included both models. It ranks third among 110 models on Vibe Code Bench v1.1 while costing about one-fifth as much as Sonnet 5.5 in that test. The gains come with heavy extended-thinking output, which can raise per-task costs and narrow the model’s advantage on long-context workloads.
Code lifts Haiku into the top three
Vibe Code Bench v1.1 asks models to build complete applications from plain-English specifications. Haiku 5.5 scored 90.44%, placing third among 110 tested models and climbing 89 positions from Haiku 4.5.
Vals reports that Haiku 5.5 passed every test for 21 of the benchmark’s 50 applications, compared with 24 for Sonnet 5.5. Haiku completed those runs at about one-fifth of Sonnet’s cost and in roughly half the time.
Across the Vals Index, a cross-benchmark ranking, Haiku 5.5 placed 16th among 45 models with a score of 54.31%. Its average cost of $2.99 per test made it the second-cheapest model in the top 16.
| Benchmark | Score | Rank |
|---|---|---|
| Vibe Code Bench v1.1 | 90.44% | 3 of 110 |
| ProofBench v1.1 | 86.00% | 10 of 48 |
| Tax Agent Bench | 62.20% | 30 of 67 |
| EMB | 59.02% | 34 of 72 |
| Finance Agent v2 | 54.14% | 24 of 76 |
| IOI | 47.28% | 27 of 42 |
| Legal Research Bench | 43.27% | 20 of 75 |
| Terminal-Bench 4.0 | 35.35% | 12 of 45 |
| Harvey’s Legal Agent Benchmark | 1.25% | 54 of 76 |
Vals’ scores reflect its prompts, tool harnesses, and run settings. Production results can vary with task design, latency limits, tool reliability, retry policies, and acceptable failure rates.
Longer reasoning traces drive the gain
On Legal Research Bench, Vals recorded an average of 59 agent steps per task for Haiku 5.5, compared with 17 for Haiku 4.5. An agent step represents a model turn or tool interaction. The newer model generated roughly 15 times as much output, and extended-thinking tokens accounted for about 80% of that output.
The same token-heavy pattern appeared against Sonnet 5.5 and Opus 5.5. On the legal research, tax, and finance benchmarks, Haiku generated more output than both larger Claude tiers, mostly through extended thinking. It used fewer agent steps than Sonnet 5.5 overall, with longer deliberation inside each model turn. Those thinking tokens contribute to metered output usage.
The 100,000-token threshold changes the economics
Haiku 5.5’s listed base rates are $0.10 per million input tokens and $0.50 per million output tokens, one-tenth of Haiku 4.5’s respective $1 and $5 rates. Requests exceeding 100,000 tokens enter a pricing tier that raises both rates fivefold.
| Context tier | Input per million tokens | Output per million tokens |
|---|---|---|
| Up to 100,000 tokens | $0.10 | $0.50 |
| Over 100,000 tokens | $0.50 | $2.50 |
Requests in the higher tier also cost $0.05 per million tokens for cache reads and $0.625 per million tokens for cache writes. Most of Vals’ Legal Research requests crossed the 100,000-token threshold.
The higher tier, combined with Haiku 5.5’s larger reasoning output, pushed its per-test cost above Haiku 4.5 on every shared benchmark. Developers evaluating agent workloads therefore need request-level estimates that include context length, thinking tokens, cache operations, retries, and tool loops.
Inside Vals’ test harness
- Context window: 1 million tokens
- Maximum output: 128,000 tokens
- Temperature: 1.0
- Sampling: Default top-p and top-k
- Compute effort: Maximum
- Fallback rate: 0.00%
- Refusal rate: 0.22%
Anthropic’s Claude model guide provides the broader model specifications, while the Vals results describe the configuration used for these benchmark runs.
Bounded workloads preserve the price advantage
| Workload | Selection signal | Primary constraint |
|---|---|---|
| Specification-to-app generation | Strong benchmark candidate | Validate framework support, correctness, and repair rates |
| Short coding and utility tasks | Low base token rates | Keep prompts and output lengths bounded |
| Long legal or financial agents | Higher per-task cost risk | Large reasoning traces and the higher context tier |
| Legal agents resembling Harvey’s benchmark | Weak benchmark result | 1.25% score and a rank of 54 among 76 models |
Cost modeling should start with representative production traces that include prompt length, answer tokens, extended-thinking tokens, cache operations, tool calls, retries, and the share of requests crossing 100,000 tokens. Vals’ data supports testing Haiku 5.5 for bounded coding workloads. Long-running agents require per-task estimates because heavy reasoning output and the higher pricing tier can consume the savings from its low base rate.