Vals AI Rebuilds Vals Index 2.1 Adding Tax and Overhauled Coding Scores
Vals Index v2.1 swaps in the harder Terminal-Bench 4 and adds a private Tax Agent Benchmark, reshuffling how frontier models get scored by GDP weight.
- Vals Index v2.1 replaces Terminal-Bench 2.1 with the harder Terminal-Bench 4 for the coding bucket.
- New Tax Agent Bench joins the index as its own GDP-weighted category alongside finance, coding, and legal.
- Claude Opus 5.5 leads at 69.69%, edging Sonnet 5.5 at 69.22% which is 37% cheaper per test.
- Terminal-Bench 4 uses mini-swe-agent with 66 all-new tasks and an 8-hour per-task time limit.
- Scoring is avg@3 pass@1 with full-verifier passes required, no partial credit, refusals count as failures.
- Full leaderboard and methodology at vals.ai/benchmarks/vals_index.
Vals Index 2.1 adds tax and rebuilds its coding score
The Vals Index has moved to version 2.1, replacing Terminal-Bench 2.1 with Terminal-Bench 4 and adding Tax Agent Bench as a separate category. Coding results now come from 66 new tasks and a different agent harness, while the composite covers U.S. corporate tax work for the first time. Earlier Terminal-Bench scores have limited comparability with the new results, and model rankings may shift even when the models remain unchanged.
Tax joins a rebuilt coding mix
Vals AI combines public and private benchmarks, then weights sector scores by their estimated shares of U.S. GDP. Before tax was added, the calculation was (8.0 × Finance + 5.6 × Coding + 1.2 × Legal) / 14.8. With tax, the structure becomes (8.0 × Finance + 5.6 × Coding + 1.2 × Legal + w × Tax) / (14.8 + w), where w is the current GDP-derived tax weight.
| Sector | Benchmarks | Composite treatment |
|---|---|---|
| Finance | Finance Agent v2 and Excel Modeling Benchmark | Combined into the 8.0-weight finance score |
| Coding | Terminal-Bench 4, Vibe Code Bench, and Code Migration | Combined into the 5.6-weight information-sector score |
| Legal | Legal Research Bench and Harvey’s HLAB | Combined into the 1.2-weight legal score |
| Tax | Tax Agent Bench | New GDP-weighted category covering research-grade U.S. corporate tax questions |
Terminal-Bench 4 starts from scratch
Terminal-Bench 4 contains 66 new tasks with no overlap with the 89 tasks in version 2.1. The harness, which mediates between the model and its tools, also changes. Version 2.1 used Terminus 2, while version 4 uses mini-swe-agent. Both versions retain Daytona’s isolated task sandboxes and pass@1 scoring, so each run records one attempt per task.
Each agent receives a task instruction and one tool: bash. On every turn, it can issue one or more shell commands, which run in a fresh subshell inside the persistent task sandbox. The agent reads the output and continues until it signals completion or reaches the eight-hour limit. No separate step or spending cap applies within that window.
Task-specific verifiers inspect the final environment or produced artifact after the agent stops. Credit requires passing the verifier’s complete test suite; a failed check yields zero for that task. This all-or-nothing design measures whether an agent completed the job, including changes that may not appear in its final response.
Vals reports avg@3 results to reduce run-to-run variance. Every model completes the full benchmark three times, each run receives a pass@1 score, and the leaderboard displays the mean of those three scores. Error bars show the standard error observed across the runs.
Anthropic takes the top four slots
Claude Opus 5.5 leads the composite at 69.69%, followed by Claude Sonnet 5.5 at 69.22%, a gap of 0.47 percentage points. Claude Fable 5.1 ranks third, Claude Opus 5 ranks fourth, and GPT-6 Astra places fifth at 66.61%. Opus 5.5 also leads Terminal-Bench 4, while Sonnet 5.5 leads Vibe Code Bench and Code Migration.
Sonnet 5.5 costs a reported $20.80 per test, compared with $32.77 for Opus 5.5. Meta’s Muse Spark 1.3 Max ranks seventh with 64.53% at $3.40 per test. GPT-5.6 Luna reaches 59.88% at $0.78 per test, trailing Opus 5.5 by 9.81 points while costing about 42 times less per test.
Code Migration uses a calibrated subset
Each benchmark column normally reproduces that benchmark’s published standalone score under the same methodology. Code Migration uses a fixed subset: 50 of the 120 CLI migration tasks plus all 10 COBOL tasks. The index weights the resulting scores 75% for CLI and 25% for COBOL, matching the standalone benchmark’s category balance.
Vals selected the subset through Monte Carlo search and reports a Spearman correlation of 0.99 with the full published results, indicating an almost identical model ranking. The reported mean absolute error is 0.8 percentage points. Those fit statistics describe the evaluated models and may change as new model families enter the benchmark.
Turning scores into a shortlist
Model-selection teams can use the index to narrow candidates by domain, test cost, and failure policy before running workload-specific evaluations. General benchmarks such as MMLU and Arena Elo do not directly test whether an agent can operate a shell, edit an Excel model, migrate code, research tax law, or cite legal authority.
- Filter by domain. Teams focused on software work can inspect Terminal-Bench 4, Vibe Code Bench, and Code Migration directly instead of relying on the composite rank.
- Compare cost with success rate. Models separated by a few accuracy points can have large differences in reported test cost, affecting production throughput and budget requirements.
- Keep blocked tasks in the denominator. Vals counts refusals and provider policy blocks as failures, preventing unavailable completions from disappearing from the score.
- Recreate deployment constraints. Finalists should run representative internal tasks under the latency, token, tool, security, and spending limits that production systems will enforce.
The rank reflects its weights
GDP weighting gives finance and insurance roughly 8% of the earlier composite, compared with 1.2% for legal services. A model with exceptional legal performance and average finance results can therefore rank below a broader model. The sector columns provide the clearer signal for specialized deployments.
Terminal-Bench also allows up to eight hours per task without separate step or spending caps. A model can pass under those conditions and still exceed a production latency or budget target. The composite works as a screening metric, while procurement decisions still require domain-specific tests under the intended operating constraints.