Google DeepMind's Gemini 4 Argon Shatters Output Limits With 1 Million Tokens
Google DeepMind's new frontier model pushes output to 1M tokens, targets long-horizon coding, enterprise workflows, and autonomous cyber defense through the Fairwind Program.
- Google DeepMind announced Gemini 4 Argon, a new frontier model for coding, enterprise, and cyber defense.
- Output token limit jumps from 64K to an industry-leading 1M tokens per generation.
- Pricing: $2/M input, $10/M output introductory; doubles after the intro period ends.
- Benchmarks: 77.9% on DeepSWE v1.1, 91.7% on LVBench, 68% on CWE-bench v1, #1 on AutomationBench.
- Rolling out via the Fairwind Program to trusted cyber defenders first, then paid API and AI Ultra.
- Internal wins include 40% quantum subroutine gains and 300+ TiB of datacenter memory freed by Argon agents.
Gemini 4 Argon raises the output ceiling to one million tokens
Google DeepMind has announced Gemini 4 Argon, a flagship model with a maximum output of one million tokens, up from 64,000. The output window determines how much the model can generate in one response; the input context window determines how much material it can read.
The larger output allowance could let one request produce an extensive code migration, audit, legal brief, or research artifact. Long generations still carry practical constraints, including latency, cost, coherence, failure recovery, and verification.
Security teams get the first look
Google is initially distributing Argon through its Fairwind Program, which gives selected cybersecurity defenders early access. The company says it is also participating in the U.S. government’s voluntary process for pre-release model access and will collect feedback before expanding availability.
Broader distribution will begin with paid API customers and Google AI Ultra subscribers. Google has promised access for developers, enterprises, and consumers, though the announcement provides no firm date for general availability.
Pricing rewards cached context
Argon’s introductory API pricing starts at $2 per million input tokens and $10 per million output tokens. Cached input receives a 95% discount, reducing the introductory rate to $0.10 per million cached tokens.
| Token type | Introductory price | Later price |
|---|---|---|
| Input | $2 per million | $4 per million |
| Cached input | $0.10 per million | $0.20 per million |
| Output | $10 per million | $20 per million |
A response that uses the full output allowance would cost $10 during the introductory period and $20 afterward, excluding input charges. Applications can set lower output limits for routine requests and reserve the full allowance for unusually large artifacts.
One call, much more output
Frontier models commonly limit responses to between 8,000 and 64,000 tokens, forcing applications to split large jobs across multiple calls. Developers then have to manage summaries, intermediate state, retries, and the assembly of partial results.
Argon’s ceiling can reduce that orchestration for workloads with large final outputs, including repository migrations, document analysis, compliance reviews, and security audits. Multi-step systems will remain useful for tool boundaries, human approvals, validation, and recovery from failed requests.
Production testing will need to measure performance near the upper end of the window. A model that remains coherent across 100,000 tokens may behave differently at one million, and a late failure can waste substantial time and token spend.
Benchmarks favor sustained work
Google’s published evaluations concentrate on software engineering, professional services, automation, video analysis, and cybersecurity. The reported results are vendor-supplied and will require independent replication.
| Benchmark | Reported result | What it measures |
|---|---|---|
| DeepSWE v1.1 | 77.9%, reported state of the art | Long-horizon software engineering on real repositories |
| Vals Index | First place | Finance, coding, legal, and tax work weighted by U.S. economic contribution |
| AutomationBench | 51.3%, first place | End-to-end execution across common business functions |
| LVBench | 91.7%, reported state of the art | Understanding and reasoning over long videos |
| CWE-bench v1 | 68%, tied for first | Finding and repairing software vulnerabilities |
Inside Google: kernels, memory, and quantum code
Google says internal teams are already using Argon agents for engineering projects that require repeated analysis, experimentation, and code modification. The examples provide more operational detail than the benchmark scores alone.
- Quantum optimization: Argon reportedly reduced the spacetime resources of bottleneck subroutines by 40% against a published baseline within minutes. Spacetime resources capture the combined qubit and runtime requirements of a quantum computation.
- Fleet memory savings: Agents analyzed profiling telemetry and applied memory optimizations across Google’s data centers, freeing more than 300 tebibytes. Google estimates the potential total savings at 500 tebibytes to one pebibyte.
- C and C++ migrations: Agents are converting codebases to Rust, ranging from core libraries such as re2 and libgav1 to more than 800,000 lines in the Fuchsia operating system’s Zircon kernel.
- Video decoding: On libgav1, agents replaced 32,000 lines of SIMD code through profile-guided experiments and compiler analysis. Google says the resulting memory-safe Rust decoder runs 2.7 times faster than the previous Rust port.
Cyber defense drives the rollout
Argon was trained for autonomous vulnerability discovery, validation, and patching, which explains the security-focused early-access program. Google says selected Fairwind participants and internal teams can use versions without the usual cyber guardrails, giving vetted defenders access to the model’s full security capabilities.
According to Google, Wiz used Argon through its Scan for Good initiative to uncover a critical flaw that exposed sensitive personal information in healthcare software used by hospitals worldwide. Previous frontier models had missed the issue.
Google also reports that Argon outperformed its predecessor, 3.8 Flash Cyber, on Wiz’s internal black-box penetration-testing benchmark. That evaluation tests attack-surface discovery, vulnerability identification, and proof-of-concept generation against live web systems without source-code access.
Guardrails depend on access
Google describes several controls for standard deployments and high-risk evaluations. The restricted release without cyber guardrails places additional weight on participant vetting, access controls, logging, incident response, and revocation policies.
- Misuse prevention: Standard deployments refuse requests involving cyberattacks or chemical, biological, radiological, and nuclear threats while preserving approved dual-use research. Red teams probe those controls, and monitoring systems inspect internal activation patterns for signs of misuse.
- Prompt-injection resistance: Google says Argon leads Gray Swan’s Indirect Prompt Injection benchmark after automated red teaming and adversarial training. Indirect prompt injection occurs when malicious instructions hidden in external content attempt to redirect an agent.
- Execution monitoring: Monitors inspect reasoning traces and actions and can stop a run. Google says it keeps monitor findings out of model training to reduce the risk that the model learns to evade oversight.
- Environment isolation: High-risk training and evaluation runs take place in sandboxed environments that are isolated and sealed before execution.
API details still needed
Argon’s production value will depend on implementation details beyond the headline token limit. Developers will need complete documentation for the following areas:
- Maximum input context and how it interacts with the one-million-token output allowance
- Streaming behavior, timeouts, cancellation, continuation, and retry semantics
- Structured output, tool calling, stop controls, and deterministic generation options
- Rate limits, concurrency quotas, service-level commitments, and regional availability
- Data retention, training-data policies, audit logs, and enterprise access controls
- Quality, latency, and failure rates across progressively longer generations
Where architecture can simplify
A million-token response ceiling reduces one source of fragmentation in coding agents, retrieval-assisted systems, document workflows, and security automation. Applications may be able to preserve more working state inside a single generation and avoid errors introduced when dozens of partial outputs are summarized and reassembled.
Argon’s practical impact will depend on reliability near the limit, API behavior, and the controls surrounding its cybersecurity capabilities. The model’s announced pricing and restricted rollout give early users a way to test those questions before broader access begins.