model release

Google DeepMind Ships Gemini 3.8 Flash and a Cybersecurity Variant, Third Flash Release in Six Weeks

TL;DR

Google DeepMind released Gemini 3.8 Flash and a specialized cybersecurity variant, Gemini 3.8 Flash Cyber, its third Flash-tier launch in six weeks. Pricing stays at $0.75 per million input tokens and $3.75 per million output tokens, matching the prior 3.7 Flash release.

3 min read
0

Google DeepMind released Gemini 3.8 Flash and a specialized variant, Gemini 3.8 Flash Cyber, on September 2, 2026 — the company's third Flash-tier model launch in six weeks, following Gemini 3.7 Flash three weeks earlier.

Gemini 3.8 Flash is priced identically to its predecessor: $0.75 per million input tokens and $3.75 per million output tokens. Google DeepMind describes it as its "most intelligent workhorse model," with gains concentrated in software engineering, agentic task execution, and multi-step reasoning.

Benchmark claims

According to Google DeepMind, Gemini 3.8 Flash scores 54.9% on HLE-Verified, a benchmark spanning STEM, humanities, and professional reasoning tasks. The company also claims the model outperforms most larger frontier models on DeepSWE v1.1, a long-horizon software engineering benchmark, and leads on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark — both domain-specific evaluations for financial and legal reasoning.

Google DeepMind attributes the gains to the model "working harder": executing additional reasoning steps and calling tools iteratively at higher effort levels, which can increase token consumption. Developers seeking lower compute overhead can select reduced effort levels or continue using Gemini 3.7 Flash, which remains supported.

Gemini 3.8 Flash Cyber

The cybersecurity variant is restricted to a new "Fairwind Program" covering government defenders, critical infrastructure operators, and software maintainers — not general API access. Google DeepMind claims the model achieves frontier-level performance on CyberGym, an industry benchmark for autonomous vulnerability discovery, and exceeds a 70% success rate on an internal benchmark spanning 20 programming languages.

On CWE-Bench, an external patching benchmark run by Collinear, the company reports a pass@1 of 47.2%, close to a competing frontier model's 47.8%, at what it describes as significantly lower cost. Google DeepMind also cites internal deployment results: its Chrome Security team reportedly saw 2.6 times more correct patches than "the best commercial models," security firm Wiz reported 7.5–9.7% higher recall at 2.3–5.2x lower cost on penetration-testing benchmarks, and Google's Cloud Vulnerability Research team says it found a critical vulnerability in under two hours using the model.

Safety measures

Both models ship with mitigations against CBRN and cyber-offense misuse under Google's Frontier Safety Framework, with Gemini 3.8 Flash Cyber operating under more permissive cyber-defense-specific controls limited to vetted users. Google DeepMind also claims improved prompt-injection robustness as measured by third-party evaluator Gray Swan.

Availability

Gemini 3.8 Flash is available now via the Gemini API, Google AI Studio, Android Studio, Google Antigravity, and Stitch, plus Gemini Enterprise and consumer surfaces including the Gemini app, AI Mode in Google Search, and Gemini in Google Sheets. Gemini 3.8 Flash Cyber access requires application through the Fairwind Program.

What this means

Three Flash-tier releases in six weeks signals Google DeepMind is prioritizing rapid iteration over infrequent, large jumps — betting that continuous small gains at fixed low pricing outcompete slower frontier-model cycles. Holding price steady while claiming performance parity with costlier models is a direct pricing pressure play against OpenAI and Anthropic's flagship tiers. The Cyber variant's gated release, rather than open API access, reflects a cautious dual-use posture: Google is willing to claim frontier-level vulnerability discovery capability publicly while restricting who can actually use it, an approach likely to become standard for offensive-adjacent security models industry-wide.

Related Articles

model release

OpenAI's Astra Model Aces Cybersecurity Benchmark, Found Two Zero-Day Exploits Unassisted

OpenAI has disclosed new details on Astra, a forthcoming model the company says is the first to cross its 'critical cybersecurity threshold.' According to OpenAI, Astra scored a perfect result on ExploitBench and discovered two zero-day vulnerabilities in internal testing without human guidance.

model release

OpenAI Says Upcoming Astra Model Is First to Cross 'Critical' Cybersecurity Risk Threshold

OpenAI says its upcoming Astra model is the first to cross its 'Critical' cybersecurity capability threshold, meaning it can discover and exploit unknown vulnerabilities without step-by-step human guidance. The company plans to release Astra soon but will restrict its advanced cyber capabilities to a vetted coalition of organizations.

model release

Anthropic's Claude Fable 5.1 Launches on Amazon Bedrock and Claude Platform on AWS

Anthropic's Claude Fable 5.1 is now live on Amazon Bedrock and Claude Platform on AWS, improving on Fable 5 in reasoning, agentic coding, and long multi-step tasks. The model ships with new Enterprise Frontier Safeguards allowing zero data retention for eligible customers through December 2026.

model release

Anthropic Releases Claude Fable 5.1, Cuts Agentic Workload Pricing Up to 45%

Anthropic has released Claude Fable 5.1, an upgrade to its top-tier Fable 5 model launched in June, alongside a restricted-access sibling called Mythos 5.1. The company claims the new model matches or beats Fable 5's performance while cutting costs by up to 45% on agentic workloads through reduced cache-read pricing.

Comments

Loading...