
The Gemini 4 Argon model is designed to handle complex, multi-step workflows while producing up to 1 million tokens in a single output, a substantial increase from the 64,000-token limit of previous Gemini generations.
Google DeepMind announced Gemini 4 Argon on September 30, positioning it as a model built for tasks that require sustained reasoning rather than short, single-step responses. The company said Argon is already being used internally for coding, research, engineering, and quantum computing, while its initial external rollout is limited to trusted cyber defenders through the Fairwind Program.

Google DeepMind announced Gemini 4 Argon, a frontier model for coding, enterprise workflows, and cybersecurity, initially rolling out to trusted testers. Source: @GoogleDeepMind via X
The launch puts Google’s latest AI model into direct competition with frontier systems, including OpenAI’s GPT-6 Astra and Anthropic’s Claude models. Google’s published benchmark comparisons show Argon leading or tying for the top result in 14 of 19 evaluations, although rival models score higher on several tests.
Gemini 4 Argon Raises the Output Limit to 1 Million Tokens
The most notable technical change is Gemini 4 Argon’s 1 million-token output limit. Google says the Gemini 4 Argon model can generate hundreds of thousands of tokens within a single reasoning trajectory, giving it more room to work through lengthy and interconnected problems.

Google introduced Gemini 4 Argon, offering frontier performance for software engineering, knowledge work, and cybersecurity with a 1M-token output limit. Source: Google via X
The company says the expanded output capacity is intended for workloads such as large-scale code migrations, extended software engineering tasks, financial research, legal work, and other enterprise processes that may require numerous steps before reaching a final result. Google describes this as an industry-leading output limit, rising from 64,000 tokens in the previous generation.
Benchmark results published by Google show Argon scoring 68.9% on the Vals Index and 51.3% on AutomationBench. It also recorded 77.9% on DeepSWE v1.1, 91.9% on Vibe Code Bench, and 88.8% on LABBench 2.
The model also performed strongly on long-context evaluations. Google reported a 99.7% score on GraphWalks tasks using up to 128,000 tokens and 84.2% on tests ranging from 256,000 tokens to 1 million tokens. These evaluations are intended to measure how well models maintain performance as the amount of information they must process increases.
The benchmark comparison with GPT-6 Astra is not uniformly in Argon’s favor. Astra posts higher scores on evaluations such as FrontierSWE v2, Terminal-Bench Science, and the OSWorld-2.0 offline subset, while Claude Opus 5.5 leads on some other coding and engineering tests. This indicates that performance varies by workload rather than one model dominating every category.
How Google Is Already Using Argon Internally
Google says Gemini 4 Argon is already being used across its engineering and infrastructure operations. Thousands of Google employees are reportedly using the model for specialized coding, research, and other technical workflows.
One application involves migrating C and C++ code to Rust. Google said Argon agents are working across projects ranging from core libraries to the Fuchsia Zircon kernel, which contains more than 800,000 lines of code. The company said large-scale rewrites are subject to automated and manual auditing, emulation testing, and additional review before production deployment.

Sundar Pichai announced Gemini 4 Argon, Google’s next frontier AI model for complex workflows, cybersecurity, software engineering, and quantum computing. Source: Sundar Pichai via X
Google also highlighted work involving libgav1, its open-source video decoder. According to the company, Argon agents replaced about 32,000 lines of performance-critical SIMD code in an existing Rust implementation. The resulting decoder runs 2.7 times faster than the earlier Rust version while producing identical video output.
Another internal use involves data center efficiency. Google said a team of Argon agents analyzed fleet-wide profiling data and identified memory optimizations that freed more than 300 tebibytes of memory after deployment. The company estimates that total savings could reach between 500 TiB and 1 pebibyte.
Argon has also been used in Google’s quantum-computing research. In one example, the company said the model helped optimize a quantum algorithmic workload and beat a published baseline by 40% in minutes. These figures are Google’s own reported results and have not yet been independently reproduced across all of the company’s internal workloads.
Cybersecurity is another major focus of Gemini 4 Argon. Google says the model can autonomously find, validate, and patch critical software vulnerabilities. Through Wiz’s Scan for Good initiative, Argon reportedly identified a critical vulnerability that could expose sensitive personal information in healthcare software used by hospitals worldwide. Google said earlier frontier models had missed the vulnerability.
On CWE-bench v1, a cybersecurity benchmark focused on vulnerability remediation, Argon tied for first place with a score of 68%.
Because of these capabilities, Google is taking a phased approach to Gemini 4 Argon. The company is initially providing access to trusted cyber defenders through its Fairwind Program while participating in the U.S. government’s voluntary pre-release model access process.
Google said early testers will help evaluate the model and refine its safeguards before broader availability. The company plans to expand access to developers, enterprises, and consumers, but it has not provided a specific public release date.
Once broader access begins, paid API customers and Google AI Ultra subscribers are expected to be among the first groups able to use the model. Google has also announced introductory API pricing of $2 per million input tokens and $10 per million output tokens, with cached input priced at a 95% discount from the standard input rate.
The restricted launch reflects the dual-use nature of increasingly capable AI systems. Gemini 4 Argon is designed to automate sophisticated technical work, but the same capabilities can create additional security risks if deployed without adequate controls. Google’s decision to begin with vetted cyber defenders allows the company to gather real-world feedback while continuing to develop safeguards ahead of a wider release.





