model release

Google Launches Gemini 4 Argon With 1M Output Tokens, Ties GPT-6 Astra on Key Benchmark

TL;DR

Google has released Gemini 4 Argon, its first frontier model in over seven months, featuring a 1 million output token limit — an industry first. Independent benchmarks show it matching OpenAI's GPT-6 Astra but trailing Anthropic's Claude Opus 5.5.

4 min read
0

Google has released Gemini 4 Argon, a new flagship model that closes the performance gap with OpenAI and Anthropic without establishing a clear lead. The model arrives more than seven months after Gemini 3.1 Pro, following a development cycle in which Google skipped its previously announced Gemini 3.5 model entirely.

Key specs

Argon accepts text, images, video, and audio as input and outputs text only. The input context window remains at 1 million tokens, but Google has raised the output limit from 64,000 to 1 million tokens — a change the company calls an industry first. To support responses of this length, Google added a "Long Decode Continuation" feature to the Gemini API that pauses long generations and resumes them via follow-up requests, avoiding timeouts during extended reasoning.

Pricing starts at an introductory rate of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 once the promotional period ends. Cached input tokens cost 95% less than standard input pricing — roughly $0.10 per million — compared to a 90% discount on the earlier Gemini 3.8 Flash.

Benchmark results: a tie, not a lead

According to independent testing from Artificial Analysis, Gemini 4 Argon at its "High" reasoning setting scores 53 on the Intelligence Index, tying OpenAI's GPT-6 Astra (max) and Anthropic's Claude Fable 5.1, and edging out GPT-6.1 Sol (max) by one point. Anthropic's Claude Opus 5.5 leads the field at 58, with Claude Sonnet 5.5 at 56. Argon's score marks a 23-point jump over Gemini 3.1 Pro Preview.

The efficiency picture is less favorable: Argon consumes an average of 62,000 output tokens per Intelligence Index task, more than double GPT-6 Astra's 27,000. At the current promotional price, one task costs $1.99 — 60% of GPT-6 Astra's $3.26 — but that advantage evaporates once regular pricing kicks in, pushing the cost to $3.98, about 20% above GPT-6 Astra.

On agentic benchmarks, a historically weak area for Gemini models, Argon shows marked improvement. It leads AutomationBench-AA at 77.5%, six points ahead of Claude Sonnet 5.5 (max). On Terminal Bench 4 it scores 57%, a 53-point jump from Gemini 3.1 Pro Preview, though it still trails Claude Sonnet 5.5 (64%), Claude Opus 5.5 (60%), and GPT-6 Astra (59%).

On the AA-Omniscience benchmark, which measures factual accuracy and honest handling of knowledge gaps, Argon posts a hallucination rate of just 15%, far below GPT-6 Astra's 51% and GPT-6.1 Sol's 54%. However, its raw accuracy is only 50% — five points below Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra's 63%. Its overall AA-Omniscience score of 42 is roughly even with GPT-6 Astra (43) and GPT-6.1 Sol (42).

Google's internal benchmarks show Argon leading by wider margins, and the model tops the Vals Index at 68.9%, ranking in the top five on 20 of 22 tested categories — the first Gemini model to top that index. On LMArena's Text Arena, Argon (High) takes first place with 1,525 points, 20 ahead of Claude Opus 4.6. It ranks first in coding, instruction following, creative writing, and professional-field queries in English, Chinese, and Russian. In Code Arena's WebDev category, however, it lands only eighth at 1,679 points — a 96-point improvement over Gemini 3.8 Flash but still short of the top ranks.

Restricted rollout

Argon is not yet broadly available. Google is first giving access to a group of "trusted cyber defenders" under its Fairwind program, along with internal teams, without cyber-related guardrails. The company says it is also part of a US government voluntary program granting agencies early access to new models. Google plans to open Argon to paying API customers and Google AI Ultra subscribers next, though it has not set a date beyond "as soon as possible."

What this means

Argon puts Google back in contention among the top three frontier labs after a rocky stretch that saw Gemini 3.5 scrapped outright. But the independent numbers tell a mixed story: Argon matches OpenAI's GPT-6 Astra on general intelligence and wins on hallucination rate and human-preference writing tests, while still trailing Anthropic's Opus 5.5 on raw capability and lagging behind Claude and OpenAI models on agentic coding tasks. The headline 1-million output token limit is a genuine technical differentiator, but Argon's heavier token consumption per task offsets much of its price advantage once the introductory discount expires. For now, Google's strongest claim is cost-efficiency and low hallucination, not outright superiority — and with access still limited to select testers, most developers will have to wait to judge for themselves.

Related Articles

model release

Google Launches Gemini 4 Argon, Restricts Initial Access to 'Trusted Cyber Defenders'

Google announced Gemini 4 Argon, a new frontier model it says excels at software engineering, enterprise knowledge work, and cybersecurity defense. The company is initially limiting access to select cybersecurity partners while it strengthens safety measures against misuse.

model release

Google DeepMind Releases Gemini 4 Argon, Expands Output Limit to 1M Tokens

Google DeepMind has released Gemini 4 Argon, a frontier model built for long-horizon reasoning with an industry-leading 1 million output token limit. The model is rolling out first to trusted cyber defenders through Google's Fairwind Program, with pricing set at $2 per million input tokens and $10 per million output tokens.

model release

Google Launches Gemini 4 Argon, Claims Top Marks in Coding and Cybersecurity Benchmarks

Alphabet launched Gemini 4 Argon on Wednesday in a phased rollout starting with trusted cybersecurity partners. Google claims the model sets a new record in real-world software engineering and ties for first place on cybersecurity benchmarks against GPT-6 Astra and Grok 4.7.

model release

Google Announces Gemini 4 Argon, Its New Frontier Model With 1M Output Tokens

Google has announced Gemini 4 Argon as its new frontier model, featuring a 1M output token limit (up from 64K) and claimed leads on coding, cybersecurity, and automation benchmarks. The model is rolling out first to Google AI Ultra subscribers and paid API customers.

Comments

Loading...