Google's Gemini 4 Argon ranks third on Artificial Analysis index, rolls out first to cyber partners
Google unveiled Gemini 4 Argon, its new flagship model, which sits behind only Claude Opus 5.5 and Claude Sonnet 5.5 on the Artificial Analysis Intelligence Index. Access begins with trusted cybersecurity partners, with no date given for wider availability.
Google unveiled Gemini 4 Argon this week, and the model ranks third on the Artificial Analysis Intelligence Index, behind only Claude Opus 5.5 and Claude Sonnet 5.5, according to CNBC's reporting. Access is limited to trusted cybersecurity partners at first, and Google has not said when the model will be more widely available.
What is confirmed
- Benchmark position: The Artificial Analysis Intelligence Index, a composite benchmark, places Gemini 4 behind only Anthropic's Claude Opus 5.5 and Claude Sonnet 5.5. The composite score itself was not published in the source material.
- Rollout: Google is launching in phases, starting with trusted cybersecurity partners. It is working with the U.S. government on pre-release safety evaluations.
- No wider timeline: According to CNBC, Google has given no indication of when the model will reach general availability.
What Google claims
Google says Argon outperforms top OpenAI and Anthropic offerings on some benchmarks. The specific benchmarks and scores were not detailed in the source, so this claim cannot be independently assessed yet.
Google also says it is already using the model internally to optimize memory in its data centers. The company claims this has freed up hundreds of terabytes of memory without buying additional hardware. Quantum computing researchers have also used the model, according to Google's release.
Tulsee Doshi, Google's Gemini model product lead, told CNBC that the phased approach "gives us more confidence" and puts a model "strong in cyber defense in the hands of defenders as soon as possible."
Not yet disclosed
- Pricing per 1M tokens: not yet disclosed
- Context window: not yet disclosed
- Parameter count and training cutoff: not yet disclosed
- General availability date: not yet disclosed
- Individual benchmark scores: not included in the source reporting
Analyst reaction
Tim Law, IDC's director of research for AI, said Argon "shows advanced reasoning on critical tasks, according to benchmarks," citing legal reasoning, finance and enterprise knowledge work, including long-running tasks.
Lian Jye Su, chief analyst at Omdia, said Google was "late to the cybersecurity and coding game" but that Gemini 4 takes it to the frontier, particularly in cybersecurity. Su also argued that security incidents involving AI systems from OpenAI and Anthropic give Google an opening to position itself as the trusted provider for secure AI operations. That is Su's characterization, not an independently established finding.
Nick Patience, AI lead at the Futurum Group, said Argon makes Google competitive again but does not make it the "leader."
Context
Gemini 3, launched toward the end of 2025, put Google in the frontier conversation. For much of 2026, however, Anthropic and OpenAI have set the pace, per CNBC. Argon is the first flagship under Koray Kavukcuoglu, who succeeded Demis Hassabis as head of DeepMind in August.
What this means
A third-place showing on one composite index puts Google back among the top tier. It does not put Google at the top, and the two models ahead of it are both Anthropic's. A composite also hides how a model performs on specific tasks, so the missing per-benchmark scores matter.
The cybersecurity-first rollout is a deliberate strategic choice. It lets Google pursue the safety and trust positioning that analysts highlight, and it limits exposure while government evaluations run. It also means builders cannot test Argon on their own workloads, so the leaderboard position is the only external signal for now.
The open questions are the ones that decide enterprise adoption: pricing, context length, latency and general availability. Until Google publishes them, Argon's competitiveness is established on a benchmark leaderboard and in Google's own internal use cases, not in production deployments.
Related Articles
Anthropic: Zhipu's Open-Weight GLM-5.3 Nearly Matches Claude Mythos Preview at Building Cyber Exploits
Anthropic's Frontier Red Team reports that Zhipu AI's open-weight GLM-5.3 comes close to Claude Mythos Preview on cyber exploit benchmarks, scoring 50/410 vs 56/410 on ExploitBench. Unlike Mythos Preview, GLM-5.3 shipped without effective safeguards and can be jailbroken with simple prompting tricks or abliteration.
Anthropic Threat Report: Claude Used for Missile Software, Mass Surveillance, and Systematic Theft by Chinese AI Labs
Anthropic's latest threat intelligence report covers December 2025 through August 2026, documenting Claude's misuse in espionage, weapons development, and nationwide surveillance operations. The report also details how seven Chinese AI labs ran covert networks—some routing their own customers' requests through Claude—to extract training data at industrial scale.
OpenAI publishes startup guide to choosing and deploying GPT-6 models, with reasoning-effort tuning
OpenAI has published a practical guide for startups building on the GPT-6 family. It covers model selection, reasoning effort, prompts and skills, tool coordination, and production workflows. The available summary discloses no pricing, context window, or benchmark figures.
Ramp AI Index: US business AI spending falls while usage rises about 50% from July peak
US companies are spending less on AI even as usage hit a record high at the end of September, according to the latest Ramp AI Index. Ramp economist Ara Kharazian attributes the drop almost entirely to price competition between OpenAI and Anthropic. In the last week of September, Anthropic took 51% of token spending and OpenAI 44.5%.
Comments
Loading...