Vals AI Crowns Claude Opus 5.5 King While MiMo Costs 100x Less

Five frontier model drops in a single week reshuffled the Vals Index leaderboard, with Claude Opus 5.5 taking the top spot and Xiaomi's MiMo closing the open-weight gap.

·
·
Vals AI Crowns Claude Opus 5.5 King While MiMo Costs 100x Less
  • Claude Opus 5.5 takes #1 on Vals Index at 69.69%, leading six benchmarks including Terminal-Bench 4.0 and ProofBench.
  • Opus 5.5 beat the human baseline on the LM Training RSI task, pulling forward Vals' RSI timeline by a year.
  • GPT-6 Sol and Luna cut cost per test by ~47% each with roughly flat accuracy versus 5.6 predecessors.
  • MiMo V2.6 Flash hits 59.6% at $0.20 per test, about 1% of Opus 5.5 pricing, and takes #1 on CyberBench.
  • Grok 4.7 climbs to #12 with big coding gains but 2.6x the cost of Grok 4.6 and a 25% ProofBench drop.
  • 18 of the top 20 Vals Index models shipped in the last three months, with gains concentrated in coding and agentic work.

Claude Opus 5.5 leads Vals as GPT-6 and MiMo cut costs

Vals AI evaluated six models released in one week through its Vals Index. The weighted aggregate combines finance, coding, and legal benchmarks designed to approximate paid work. Several included agentic tests require models to use tools and complete multistep tasks.

The results put Claude Opus 5.5 at the top, show OpenAI reducing inference costs, and place Xiaomi’s MiMo models near the leading proprietary systems for a fraction of the price. Aggregate rankings still obscure domain regressions, model-routing behavior, and large differences in token use.

Results at a glance
Model Vals Index Standing Cost signal
Claude Opus 5.5 69.69% #1 of 63 $22.30 per test
GPT-6 Sol 62.6% #8 $7.56 per test
Grok 4.7 60.2% #12 About 2.6× Grok 4.6
MiMo V2.6 Flash 59.6% CyberBench leader $0.20 per test
MiMo V2.6 Pro 59.5% #2 open-weight model $0.39 per test
GPT-6 Luna 58.5% #20 $0.42 per test

Opus 5.5 wins on agents and spends more

Claude Opus 5.5 finished ahead of Claude Fable 5.1 at 68.83%, Claude Opus 5 at 67.21%, and GPT-6 Astra at 66.61%. Vals reports first place on six leaderboards, with published highlights including Terminal-Bench 4.0 at 61.62%, MedScribe at 91.43%, the Vals RSI Index at 37.13%, and a perfect 100% on ProofBench v1.1.

On the Recursive Self-Improvement Index’s language-model training task, Opus 5.5 trained a small model within a 24-hour budget and exceeded the published human baseline. Vals consequently advanced its projected date for full recursive self-improvement by one month from August 2027. That projection is an extrapolation from a constrained benchmark task rather than a direct demonstration of unrestricted self-improvement.

ProgramBench produced one of the clearest gains: Opus 5.5 fully resolved 18.5% of tasks, compared with 7% for Fable 5.1, 5.5% for GPT-6 Astra, and 3% for Opus 5. Its runs finished in 2.4 hours and cost $42.66 per task.

Compared with Opus 5, the newer model raised average cost per test from $18.81 to $22.30. It also regressed on MedCode, Public Benefits, Legal Research, Tax Agent Bench, and HLAB, with losses concentrated in retrieval, source fidelity, and rubric compliance. Its 3.75% score on Harvey’s Legal Agent Benchmark placed it #31 of 64.

Routing and token use complicate deployment

  • Anthropic’s server-side safety controls reroute most cybersecurity requests to Claude Opus 4.8. Biology tasks fall back to Claude Opus 5 unless the account is verified. Treating fallback-assisted tasks as failures lowers the reported SRE Bench score from 33.59% to 5.34%.
  • Artificial Analysis measured roughly 119,000 output tokens per Intelligence Index task, compared with about 73,000 for Opus 5, 78,000 for Fable 5.1, and 27,000 for GPT-6 Astra. Thinking tokens are billed as output tokens, so always-on adaptive reasoning contributes directly to the higher test cost.

GPT-6 cuts prices as scores hold near prior levels

GPT-6 Sol scored one point below GPT-5.6 Sol’s 63.7%, while reducing cost per test by 47% from $14.21 to $7.56. It improved by 7.3% on Vibe Code Bench, 18.5% on Terminal-Bench Science, and 6.5% on Terminal-Bench 4.0. Legal Research and Tax Agent scores declined.

GPT-6 Luna reduced cost by 46% to $0.42 per test. It posted small gains on Terminal-Bench, MysteryMechanism, and Vibe Code Bench, alongside a small regression on Finance Agent. The two releases shift OpenAI’s price-performance curve through lower inference costs rather than higher aggregate accuracy.

MiMo reaches 59.6% for 20 cents

MiMo V2.6 Flash reached 59.6% on the Vals Index for $0.20 per test, about 1% of the Opus 5.5 cost. It also led CyberBench with a score of 75.4%.

MiMo V2.6 Pro scored 59.5% and rose 27 positions from V2.5 Pro’s #44 finish. Its largest reported benchmark gains were:

  • Vibe Code Bench: 51.1%
  • ProofBench: 48%
  • Legal Research: 31.2%, reaching #11

The Pro release now ranks second among open-weight models at $0.39 per test. Its combination of price, accessible weights, and improved coding scores makes it a practical candidate for throughput-heavy pipelines, subject to testing against the target workload.

Grok 4.7 gains in coding, slips in proofs

Grok 4.7 rose from Grok 4.6’s #19 position to #12. Its Vibe Code Bench score increased by almost 10 points to 86.2%, Terminal-Bench 4.0 gained 11 points, and Harvey’s Legal Agent Benchmark placed it #7. Cost per test increased to about 2.6 times that of Grok 4.6, while ProofBench fell by 25%.

New releases crowd the top 20

Vals reports that 18 of the top 20 models shipped within the past three months. Most recent gains cluster in coding and agentic tool use, with less movement on knowledge and general reasoning tests.

The aggregate gap between the leading proprietary model and the top MiMo result is about 10 percentage points. The corresponding test-cost gap exceeds 100×, creating room for systems that route routine work to cheaper models and reserve premium models for difficult tasks.

A workload-first shortlist

Workload Candidate to test Reason
Agentic coding and terminal work Claude Opus 5.5 Leading ProgramBench, Terminal-Bench, and aggregate results
Legal or tax pipelines Claude Opus 5 or Fable 5.1 Opus 5.5 regressed on legal research, tax, retrieval, and rubric compliance
Low-cost, high-throughput inference MiMo V2.6 Flash 59.6% aggregate score at $0.20 per test
Open-weight deployment MiMo V2.6 Pro #2 open-weight ranking with strong coding gains
Lower-cost OpenAI workloads GPT-6 Sol or Luna Substantial price reductions with aggregate scores near their predecessors

Production selection should pair the Vals shortlist with workload-specific evaluation, including latency, token consumption, fallback routing, retrieval accuracy, and rubric compliance. Those factors can reverse the apparent advantage shown by a single aggregate score.

Trending
  • No trending articles

Comments

avatar

Next Reads