Google Announces Gemini 4 Argon, Its New Frontier Model With 1M Output Tokens
Google has announced Gemini 4 Argon as its new frontier model, featuring a 1M output token limit (up from 64K) and claimed leads on coding, cybersecurity, and automation benchmarks. The model is rolling out first to Google AI Ultra subscribers and paid API customers.
Google has announced Gemini 4 Argon, a new frontier model the company describes as its "next era of frontier intelligence," built for what it calls "deep reasoning across complex, long-horizon workflows." The release marks a shift in Google's naming convention and follows the cancellation of the previously planned Gemini 3.5 Pro, with development resources instead directed toward Gemini 3.8 Flash and now Argon.
The headline technical change is output capacity: Gemini 4 Argon's output token limit jumps to 1 million tokens, up from 64,000 in prior models. According to Google, this expanded headroom lets the model "think deeply" and generate hundreds of thousands of tokens in a single trajectory, aimed at solving complex problems in one pass rather than across multiple turns.
Benchmark claims
Google reports a score of 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%), according to the company's own figures. On CWE-bench v1, which evaluates a model's ability to remediate security vulnerabilities, Google says Argon ties for first place at 68%. The company also claims a #1 ranking on Zapier's AutomationBench, which measures end-to-end business-function execution, with a score of 51.3%, and state-of-the-art performance on LVBench (long video understanding) at 91.7%. These figures come from Google and have not been independently verified.
Cybersecurity and safety posture
Google says it trained Argon to be "highly capable at cybersecurity defense," with gains over Gemini 3.8 Flash Cyber. A version without cyber guardrails is being made available to trusted defenders and internal teams through Google's Fairwind Program, intended to let vetted users access the model's full defensive capabilities without restriction.
Ahead of wider release, Google says it is prioritizing several safety measures: monitoring internal model activations to detect misuse related to cyber or CBRN (chemical, biological, radiological, nuclear) risks; hardening resilience against indirect prompt injection, where Google claims leading results on Gray Swan's IPI Benchmark; sealing sandboxed environments before high-risk training or evaluation; and deploying misalignment monitoring that tracks the model's chain-of-thought and can halt execution. Google states it took care to avoid feeding monitoring findings back into training, to prevent the model from learning to evade oversight.
Internal deployment
Google says Argon is already running internal workflows, with "thousands of Googlers" using it for coding, research, and writing tasks. Cited examples include fleet-wide memory optimization across Google's data centers — reportedly freeing over 300 TiB immediately, with an estimated 500 TiB to 1 PiB in total savings — and large-scale C/C++-to-Rust migrations, including an 800,000+ line rewrite of the Fuchsia OS Zircon kernel. For the libgav1 video decoder, Google says Argon agents replaced 32,000 lines of SIMD code, producing a memory-safe Rust version that runs 2.7x faster than a prior Rust port while matching C++ performance.
Availability
Gemini 4 Argon is described as "rolling out soon," starting with Google AI Ultra subscribers and paid API customers. It has already reached trusted testers and cyber defenders via the Fairwind Program. Google has not disclosed pricing or full context window specifications.
What this means
Argon's benchmark claims, if verified independently, would put it slightly ahead of Claude Opus 5.5 and GPT-6 Astra on coding tasks — a narrow margin that underscores how tightly frontier labs are now clustered on core reasoning benchmarks. The 1M output token expansion is the more consequential engineering change: it enables single-pass agentic workflows (like the Rust migrations Google describes) that previously required chaining multiple model calls. The unguarded cybersecurity variant for vetted defenders signals Google's confidence in monitoring infrastructure, but it also raises the stakes if that access is ever misused or leaked. As with all self-reported benchmarks, independent evaluation will be the real test of Argon's claimed lead.
Related Articles
Google DeepMind Releases Gemini 4 Argon, Expands Output Limit to 1M Tokens
Google DeepMind has released Gemini 4 Argon, a frontier model built for long-horizon reasoning with an industry-leading 1 million output token limit. The model is rolling out first to trusted cyber defenders through Google's Fairwind Program, with pricing set at $2 per million input tokens and $10 per million output tokens.
Google Launches Gemini 4 Argon, Restricts Initial Access to 'Trusted Cyber Defenders'
Google announced Gemini 4 Argon, a new frontier model it says excels at software engineering, enterprise knowledge work, and cybersecurity defense. The company is initially limiting access to select cybersecurity partners while it strengthens safety measures against misuse.
Google Launches Gemini 4 Argon, Claims Top Marks in Coding and Cybersecurity Benchmarks
Alphabet launched Gemini 4 Argon on Wednesday in a phased rollout starting with trusted cybersecurity partners. Google claims the model sets a new record in real-world software engineering and ties for first place on cybersecurity benchmarks against GPT-6 Astra and Grok 4.7.
Google Releases Gemini 4 Argon to Cybersecurity Partners, Claims Wins Over GPT-6 Astra
Google has released Gemini 4 Argon, its next-generation flagship AI model, to a small group of cybersecurity partners as part of a phased rollout. The company claims the model outperforms OpenAI's GPT-6 Astra on several coding and knowledge-work benchmarks, though full specifications remain undisclosed.
Comments
Loading...