Google Rolls Out Gemini 3.8 Flash, Third Flash Update in Three Months
Google has released Gemini 3.8 Flash, the third Flash-tier update in three months, arriving just three weeks after Gemini 3.7 Flash. The model is live now in the Gemini app, Google Antigravity, and AI Studio with introductory pricing of $0.75/1M input and $3.75/1M output tokens.
Google began rolling out Gemini 3.8 Flash on September 2, 2026, just three weeks after its previous Flash release — the third update to the Flash line in three months. The model is already live for Google AI subscribers in the Gemini app, as well as in Google Antigravity and AI Studio.
Google describes 3.8 Flash as its "most intelligent Flash model," engineered specifically for long-horizon software engineering tasks, autonomous agents, and complex enterprise workflows. According to Google, the model beats its predecessor across "various benchmarks," though the company has not published specific scores alongside the announcement.
Pricing and availability
Google is offering an introductory rate of $0.75 per 1 million input tokens and $3.75 per 1 million output tokens, matching the pricing structure used for the prior Flash release. That introductory pricing holds through December 31, 2026, after which standard rates are expected to apply — though Google has not disclosed what those will be.
Knowledge cutoff stays split
As with the previous Flash model, Gemini 3.8 Flash carries an inconsistent training cutoff. Google states the model's knowledge extends to March 2026 "for some domains," while users may find the model's knowledge limited to January 2025 in others. This split-cutoff disclosure has now appeared in back-to-back Flash releases, suggesting Google is continuing to blend fresher fine-tuning data with an older base training set rather than fully retraining the model from scratch each cycle.
No changes to context window size have been disclosed with this release, and Google has not published parameter counts, as is standard practice for the Gemini line.
What this means
Three Flash-tier releases in three months signals Google is treating Flash as its highest-velocity iteration lane — a place to ship incremental gains in coding and agentic performance without the fanfare of a Pro-tier launch. That cadence is unusual even by current industry standards, where most labs space major or minor model updates weeks to months apart, and it raises a real question: how much of this is genuine model improvement versus rapid A/B testing of variants against benchmarks and pricing tiers.
The split knowledge cutoff — March 2026 for some domains, January 2025 for others — is the detail worth watching. It suggests Google is not doing full retrains between releases, instead patching or fine-tuning a shared base model to hit agentic and coding benchmarks faster. For developers building on Flash, that means checking task-specific outputs for recency gaps rather than assuming a uniform cutoff.
The held-over introductory pricing through year-end keeps Flash competitive on cost against Anthropic's and OpenAI's comparable fast-tier models, but the real signal is velocity: if Google sustains a three-week release cycle, Flash becomes less a stable product and more a continuously tuned service, which changes how enterprises should think about version pinning and regression testing in production.
Related Articles
Google Adds Agentic Video Understanding to Gemini, Cutting Token Use by Up to 88%
Google DeepMind has launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, letting models dynamically scan video segments instead of processing at a fixed frame rate. The company claims the feature cuts token consumption by up to 88%, reduces costs by up to 66%, and improves accuracy by up to 7%.
Google Adds Agent-Based Video Analysis to Gemini Flash, Cutting Token Usage by Up to 88 Percent
Google is rolling out agent-based video analysis for Gemini Flash models that dynamically searches footage instead of scanning frame by frame. Google claims the approach cuts token usage by up to 88 percent and costs by 66 percent while improving accuracy, with no added API fee.
Anthropic Releases Claude Fable 5.1 and Mythos 5.1, Cuts Cache Pricing 75% But Output Tokens Jump 70%
Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, claiming the top spot on Artificial Analysis's Intelligence Index at 66. Cache-read pricing dropped 75% to $0.25 per million tokens, but a 1.7x increase in output token usage pushes net per-task cost up 20%.
Anthropic to Cut Claude Code Weekly Limits by 17% Despite Calling It a 25% Increase
Anthropic will permanently raise Claude Code's baseline weekly usage limits by 25% starting September 14. Because this replaces a temporary 50% boost currently active, users will actually end up with about 17% less capacity than they have today.
Comments
Loading...