Google's Gemini 4 Argon Tops Vals Index at 68.9% Using Fewer Tokens
Google's new flagship Gemini 4 Argon claims the top slot on the Vals Index at 68.9%, beating Sonnet 5.5, Opus 5, and GPT-6 while using far fewer tokens.
- Gemini 4 Argon takes #1 on the Vals Index at 68.9%, Google's first top finish there.
- Priced at $4 in / $20 out per million tokens, $2/$10 during introductory pricing.
- Uses ~25% of Sonnet 5.5's output tokens and ~33% of its input tokens per task.
- Wins Finance Agent v2, Harvey Legal Agent, and perfect 100% on IOI 2024-2026.
- Terminal-Bench 4.0 tripled to 57.6% vs Gemini 3.8 Flash's 19.0%.
- Rolling out first to cybersecurity partners before broader release.
Gemini 4 Argon tops Vals with leaner agent runs
Google’s Gemini 4 Argon has reached No. 1 on the Vals Index with a score of 68.9%, giving Google its first outright lead on the benchmark. The index tests work across finance, law, software development, and tax using tasks designed to approximate professional workflows. Argon’s result pairs the highest aggregate score with lower token use than several close competitors, which could reduce the cost of long-running agents.
The lead costs fewer tokens
Vals reports an estimated cost of $15.68 per task and a median completion time of about 46 minutes under its test configuration.
| Metric | Result |
|---|---|
| Vals Index score | 68.9% |
| Estimated cost per task | $15.68 |
| Median completion time | About 46 minutes |
| Standard input price | $4 per million tokens |
| Standard output price | $20 per million tokens |
| Introductory input price | $2 per million tokens |
| Introductory output price | $10 per million tokens |
The 46-minute figure measures completion time for benchmark tasks that can include extended reasoning and tool calls. Interactive request latency will vary by workload. Google’s Logan Kilpatrick reported the introductory token rates, so teams calculating production costs should use the prices available to their accounts.
Vals describes Argon as moving the efficiency frontier: every cheaper model in its comparison scored lower, and Argon cost substantially less than its nearest accuracy rivals. On Vals Index tasks, it used roughly one-quarter of Claude Sonnet 5.5’s output tokens and one-third of its input tokens for comparable work.
Agent loops shrink
On Tax Agent, Argon averaged 19 turns compared with Sonnet 5.5’s 47, while producing answers more than twice as long. On Finance Agent v2, it generated about one-fifth as many tool errors. Shorter agent loops reduce tool round trips, retries, and context growth, all of which affect latency and cost in production systems.
Legal, finance, and code set the pace
Across the 22 benchmarks reported by Vals, Argon placed among the top five on 20. Its strongest results covered several distinct workloads:
| Area | Result | Developer relevance |
|---|---|---|
| Legal work | Fully completed roughly seven times as many Harvey Legal Agent tasks as Sonnet 5.5 and met nearly every grading criterion across 24 practice areas. | Supports document analysis and multistep legal workflows. |
| Finance | Ranked No. 1 on Finance Agent v2 with 65.4%, including strong results on earnings filings and disclosures. | Favors pipelines that extract and reconcile evidence from financial documents. |
| Application coding | Earned perfect results on 30 Vibe Code Bench v1.1 apps, ahead of Claude Opus 5 at 25 and GPT-6 Astra at 24, with roughly one-third fewer tool calls. | Suggests higher completion rates with less agent overhead. |
| Terminal tasks | Scored 57.6% on Terminal-Bench 4.0, up from Gemini 3.8 Flash’s 19.0%, in half the wall-clock time. | Shows a substantial gain on command-line work, although another model still leads the benchmark. |
| Competitive programming | Scored 100% on the IOI 2024, 2025, and 2026 sets, matching GPT-6 Astra. | Indicates strong algorithmic problem-solving under benchmark conditions. |
| Security | Ranked No. 1 on CyberBench’s proof-of-concept binary exploitation tasks at 70.0% and No. 2 overall behind GPT-6 Sol. | Places it 18 points ahead of Sonnet 5.5 on the overall benchmark. |
| Code migration | Raised the share of passing tests from about one-quarter with Gemini 3.8 Flash to about two-thirds. | Supports large refactoring and framework migration evaluations. |
| Site reliability engineering | Scored 44.3% on SRE Bench, 14 points ahead of Sonnet 5.5. | Provides a stronger baseline for incident and infrastructure agents. |
Four benchmarks expose gaps
- Computer use: Argon ranked seventh of eight models on CUA-bench with 4.83%, limiting the evidence for agents that operate graphical interfaces.
- Medical scribing: Its 87.4% MedScribe score placed 15th, showing that a high raw score can still trail a tightly clustered field.
- Program synthesis: A 2.5% ProgramBench score points to weak performance on formal code generation from specifications.
- Terminal work: Claude Opus 5.5 retained the Terminal-Bench 4.0 lead with 66.4%, compared with Argon’s 57.6%.
Inside the Vals run
Vals tested Argon through Google’s provider using temperature 1, default top-p and top-k settings, and high reasoning effort. Temperature, top-p, and top-k control how the model samples candidate tokens, while reasoning effort governs how much computation it devotes to solving a task. Changes to those settings can affect accuracy, token consumption, and latency.
The benchmark configuration provided a one-million-token context window and capped output at 262,000 tokens. Google’s announcement advertises continuous reasoning and generation trajectories of up to one million output tokens, up from 64,000 in the previous generation. The leaderboard therefore reflects the 262,000-token cap used by Vals rather than the model’s full advertised output ceiling.
Access starts behind a gate
Argon’s lead ends a run of Vals Index wins by flagship models from Anthropic and OpenAI. Gemini 4 also returns Google’s focus to its highest-capability tier after a year centered on faster, lower-cost Flash models.
Google is beginning a phased rollout with trusted cybersecurity partners while it evaluates the model’s safety with the U.S. government. Broader availability through the Gemini API is expected later, although Google has provided no release date.
Where Argon fits first
Teams running legal or financial document pipelines, code migrations, terminal agents, and other tool-heavy workflows have the clearest reasons to evaluate Argon. Its combination of domain accuracy, fewer agent turns, lower tool-error rates, and reduced token consumption could improve cost per successful task.
Computer-use agents, medical scribing, and formal program synthesis require separate evaluations because Argon trails on the corresponding benchmarks. A production bake-off should also measure end-to-end completion rate, tool reliability, latency, applicable token prices, rate limits, and data-governance requirements. Those results will determine whether the leaderboard advantage carries into a specific system.