Google DeepMind Releases Gemini 4 Argon, Expands Output Limit to 1M Tokens
Google DeepMind has released Gemini 4 Argon, a frontier model built for long-horizon reasoning with an industry-leading 1 million output token limit. The model is rolling out first to trusted cyber defenders through Google's Fairwind Program, with pricing set at $2 per million input tokens and $10 per million output tokens.
Google DeepMind announced Gemini 4 Argon on September 30, 2026, a frontier model designed for sustained reasoning across complex, long-horizon tasks in software engineering, enterprise knowledge work, and cybersecurity defense.
The headline technical change is output capacity: Argon's output token limit jumps to 1 million tokens, up from 64,000 tokens in prior models. According to Google, this expanded headroom lets the model "think deeply" and generate long single-trajectory responses for difficult, multi-step problems rather than breaking them into multiple calls.
Pricing and Availability
Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted 95% below standard input pricing. The model is not yet broadly available. Google is rolling it out first to "trusted cyber defenders" through its Fairwind Program and says it is participating in the U.S. government's voluntary pre-release model access process before expanding to developers, enterprises, and consumers.
Benchmark Performance
Google reports the following scores, which are company-claimed and not independently verified:
- DeepSWE v1.1 (real-world long-horizon software engineering): 77.9%, described as a new state of the art
- Vals Index (economic impact across finance, coding, legal, and tax, weighted by U.S. GDP contribution): leading position
- AutomationBench (Zapier's end-to-end business execution benchmark): 51.3%, ranked #1
- LVBench (long video understanding): 91.7%, described as state of the art
- CWE-bench v1 (vulnerability remediation): 68%, tied for first place, building on the prior 3.8 Flash Cyber model's performance on CWE-bench v0
Google did not disclose an input context window figure alongside the new 1M output limit.
Cybersecurity Focus
Argon was trained specifically for defensive cybersecurity work, according to Google, including autonomously finding, validating, and patching software vulnerabilities. For trusted defenders and internal Google teams, the company says it will release a version of Argon "without cyber guardrails" to enable full use of its vulnerability-discovery capabilities.
Google cites an early deployment through Wiz's Scan for Good initiative, in which Argon reportedly identified a critical vulnerability exposing personal health information in hospital software — a risk Google says earlier frontier models had missed. On Wiz's internal black-box penetration testing benchmark, Google claims Argon outperforms its previous 3.8 Flash Cyber model at identifying attack surfaces and producing proof-of-concept exploit evidence.
Internal Use at Google
Google says Argon is already running internal workflows at the company. Cited examples include a 40% improvement over a published baseline in quantum algorithm resource optimization, autonomous identification of memory optimizations across Google's data centers (claimed savings of over 300 TiB already deployed, with an estimated 500 TiB to 1 PiB total), and large-scale C/C++-to-Rust codebase migrations, including a 800,000+ line rewrite of the Fuchsia OS Zircon kernel. For the libgav1 video decoder, Google claims an Argon-driven Rust rewrite runs 2.7x faster than a prior Rust port while producing identical video output.
What This Means
Argon's phased, defender-first rollout signals Google is treating this release as high-risk relative to prior Gemini launches, explicitly citing its Frontier Safety Framework and CBRN misuse concerns as reasons for gating access. The pricing — $2/$10 per million tokens — sits above Gemini's typical consumer-tier pricing but remains competitive with other frontier reasoning models. The real test will come when Argon reaches general availability: internal benchmarks and company-reported wins in quantum optimization and codebase migration are notable but unverified by outside parties. Until independent evaluators can run their own tests, especially on the cybersecurity vulnerability-discovery claims, the practical gap between Argon and existing frontier models remains an open question.
Related Articles
Google Announces Gemini 4 Argon, Its New Frontier Model With 1M Output Tokens
Google has announced Gemini 4 Argon as its new frontier model, featuring a 1M output token limit (up from 64K) and claimed leads on coding, cybersecurity, and automation benchmarks. The model is rolling out first to Google AI Ultra subscribers and paid API customers.
Google Launches Gemini 4 Argon, Restricts Initial Access to 'Trusted Cyber Defenders'
Google announced Gemini 4 Argon, a new frontier model it says excels at software engineering, enterprise knowledge work, and cybersecurity defense. The company is initially limiting access to select cybersecurity partners while it strengthens safety measures against misuse.
Google Releases Gemini 4 Argon to Cybersecurity Partners, Claims Wins Over GPT-6 Astra
Google has released Gemini 4 Argon, its next-generation flagship AI model, to a small group of cybersecurity partners as part of a phased rollout. The company claims the model outperforms OpenAI's GPT-6 Astra on several coding and knowledge-work benchmarks, though full specifications remain undisclosed.
Google Launches Gemini 4 Argon, Claims Top Marks in Coding and Cybersecurity Benchmarks
Alphabet launched Gemini 4 Argon on Wednesday in a phased rollout starting with trusted cybersecurity partners. Google claims the model sets a new record in real-world software engineering and ties for first place on cybersecurity benchmarks against GPT-6 Astra and Grok 4.7.
Comments
Loading...