Google's Gemini 4 Argon Pushes AI Agent Output to One Million Tokens
Google's new frontier model targets long-horizon coding, legal, finance, and cyber defense workflows with a 1M token output ceiling.
- Google announced Gemini 4 Argon, a frontier model for long-horizon coding, enterprise, and cyber defense work.
- Output token limit expanded from 64K to an industry-leading 1M tokens for deeper single-shot reasoning.
- Introductory pricing: $2 per million input tokens, $10 per million output, 95% cached input discount.
- State of the art on DeepSWE v1.1 (77.9%), Vals Index, AutomationBench (51.3%), and LVBench (91.7%).
- Rolling out first to cyber defenders via the Fairwind Program, without cyber guardrails for trusted testers.
- Internal wins include a 2.7x faster memory-safe libgav1 decoder and 300+ TiB of freed data-center memory.
Gemini 4 Argon raises agent output limit to one million tokens
Google has announced Gemini 4 Argon, a high-end model designed for long-running work such as production coding, legal drafting, financial research, and security testing. Its defining technical change is a one-million-token output limit, which gives agents more room to plan, use tools, revise work, and complete large tasks within a single generation.
Initial access is limited to selected cyber defenders in Google’s Fairwind Program and the company’s internal teams. Google plans to expand availability after additional safety testing, beginning with paid API customers and Google AI Ultra subscribers. The announcement provides no date for broader access.
One million tokens out
Argon increases the maximum output from 64,000 tokens to one million, a roughly 15.6-fold jump. The output limit governs how much text or code the model can generate. The context window governs how much information it can read and consider.
Large software migrations, penetration tests, and document-generation jobs can exhaust an output allowance before reaching the context limit. A larger budget reduces the need to split those jobs across multiple model calls, where summaries, state transfers, and orchestration errors can disrupt the work.
The higher ceiling removes one constraint without guaranteeing a coherent million-token response. Argon’s benchmark results and internal deployments provide Google’s evidence that the model can sustain useful work over longer runs.
A gated launch, then API access
Google will introduce Argon at $2 per million input tokens and $10 per million output tokens. Prices double after the introductory period, although the company has not said when that period ends.
| Token type | Introductory price | Later price |
|---|---|---|
| Input | $2 per million | $4 per million |
| Cached input | $0.10 per million | $0.20 per million |
| Output | $10 per million | $20 per million |
A response that uses the full output allowance would cost $10 during the introductory period and $20 afterward, before input charges. Cached input receives a 95% discount, which could reduce costs for agents that repeatedly reuse large prompts, repositories, or reference documents.
Benchmarks built around finished work
Google emphasizes evaluations of long-running technical and commercial tasks. The company reports the following results:
| Benchmark | Argon result | What it measures |
|---|---|---|
| DeepSWE v1.1 | 77.9%, best reported score | Long-horizon software engineering on realistic tasks |
| Vals Index | First place | Finance, coding, legal, and tax work weighted by contribution to U.S. GDP |
| AutomationBench | 51.3%, first place | Zapier’s evaluation of end-to-end business automation |
| LVBench | 91.7%, best reported score | Understanding of long videos |
| CWE-bench v1 | 68%, tied for first | Remediation of software security vulnerabilities |
These benchmarks cover distinct tasks and scoring methods, so their percentages are not directly comparable. Together, they support Google’s focus on agents that execute extended workflows across software, security, and professional services.
Cyber defenders get the first build
Google chose security teams for the first external deployment because Argon can reportedly find, validate, and patch critical vulnerabilities with limited human intervention. Selected defenders and Google’s internal teams receive access without the standard cyber guardrails, allowing them to test the model’s full security capabilities in controlled environments.
According to Google, Wiz used Argon through its Scan for Good initiative to identify a critical vulnerability in healthcare software used by hospitals worldwide. The flaw exposed sensitive personal information and had escaped detection by previous frontier models. Google also reports gains in black-box penetration testing, where the model probes a live system without access to its source code.
Those capabilities can also support offensive activity, which explains the restricted release and additional review before general API access. Google says it is participating in the U.S. government’s voluntary process for pre-release model access while collecting feedback from early testers.
Google puts Argon to work
Google describes three internal deployments that show how the model handles extended technical jobs:
- Quantum optimization: Argon improved a published baseline for a bottleneck quantum subroutine by 40% within minutes.
- Fleet-wide memory tuning: Multiple Argon agents analyzed profiling telemetry and applied memory optimizations across Google’s data centers. Deployed changes freed more than 300 TiB, with estimated total savings between 500 TiB and 1 PiB.
- C++ to Rust migration: Argon agents are converting C and C++ codebases to Rust, from libraries containing tens of thousands of lines to more than 800,000 lines in the Fuchsia Zircon kernel. For the libgav1 video decoder, Argon replaced 32,000 lines of SIMD code with memory-safe Rust that the compiler could automatically vectorize. Google says the result produced identical video output and ran 2.7 times faster than the previous Rust port.
The libgav1 project illustrates the intended benefit of the larger output allowance: an agent can inspect profiles, run experiments, revise code, and validate behavior across a migration without repeatedly compressing its state into new sessions.
Safety controls watch the run
Google says Argon combines restricted access with monitoring designed for long autonomous executions:
- Activation monitoring: Internal systems inspect model signals for patterns associated with misuse. Google says internal and external red teams tested these techniques.
- Prompt-injection hardening: Automated red teaming and adversarial training target malicious instructions hidden in websites, documents, or other content an agent reads. Google reports that Argon leads Gray Swan’s Indirect Prompt Injection benchmark.
- Reasoning and action monitoring: Monitors inspect Argon’s reasoning traces and actions, then stop execution when they detect dangerous behavior.
Google also says it keeps findings from reasoning monitors out of the model’s training data. That separation aims to prevent training from rewarding behavior that conceals the patterns the monitors are designed to detect.
The API questions still open
Developers evaluating Argon still need details that Google has not published, including:
- The general API release date, model identifier, supported regions, and rate limits
- Latency and reliability as generations approach the one-million-token ceiling
- Streaming, checkpointing, interruption recovery, and tool-call behavior during long runs
- Output quality and error accumulation across hundreds of thousands of tokens
- Administrative controls and cyber restrictions for ordinary API customers
Argon’s practical value will depend on how well those operational details support sustained production use. The announced output limit, pricing, benchmark results, and internal deployments position the model for coding agents, enterprise research systems, document automation, and defensive security workflows once API access opens.