Google Launches Gemini 3.8 Flash, Warns It May Use More Tokens Despite Unchanged Pricing
Google released Gemini 3.8 Flash just weeks after Gemini 3.7 Flash, keeping the same per-token pricing of $0.75/$3.75 per million input/output tokens but warning it may consume more tokens overall. The model also ships with a cyber-focused variant restricted to a new government partner program called Fairwind.
Google has released Gemini 3.8 Flash, an update to its Flash model line that arrives just weeks after Gemini 3.7 Flash. The company says the new model "works harder" than its predecessor by executing more reasoning steps on complex tasks and "calling tools iteratively," but that added effort could mean higher costs for users despite unchanged per-token pricing.
Pricing stays flat, but usage may rise
Gemini 3.8 Flash carries the same introductory pricing as Gemini 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens. However, Google warns that "the model might use more tokens to maximize performance, especially at higher effort levels." Developers who want to minimize token consumption can continue using Gemini 3.7 Flash instead.
Third-party analysis firm Artificial Analysis backed this up, calling Gemini 3.8 Flash "the cheapest we've measured at this level of intelligence" on a per-task basis, while noting actual costs are up roughly 40% from Gemini 3.7 Flash. According to Artificial Analysis, that increase stems from a 30% rise in output tokens per task and more turns during agentic evaluations — meaning the model does more work per request, which drives total spend higher even though the per-token rate hasn't changed.
Benchmark claims
According to Google, Gemini 3.8 Flash delivers "significant improvements" for software engineering and autonomous agent tasks. The company claims it outperforms both Gemini 3.7 Flash and competing frontier models — including Anthropic's recently updated Fable 5 — on the DeepSWE v1.1 software engineering benchmark. Google also claims wins on the Vals Finance Agent V2 benchmark and Harvey's Legal Agent benchmark, though none of these scores were disclosed publicly in specific numeric form.
Early reactions from outside Google echoed the performance claims. Aigora.ai CEO John Ennis said the model offers "Opus 5 coding quality but at a fraction of the cost and super fast," specifically praising it for tasks like generating Remotion videos.
Safeguards and a new cyber-focused variant
Google says Gemini 3.8 Flash "ships with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense." Alongside the main release, Google launched Gemini 3.8 Flash Cyber, a variant restricted to a new access program called Fairwind. Fairwind is limited to governments and "trusted partners," with roughly 650 members including CrowdStrike and the Center for Internet Security. Members get access to 3.8 Flash Cyber and Google's CodeMender agent, which the company says can "autonomously find and fix vulnerabilities" in critical infrastructure and public services.
Gemini 3.8 Flash is available now to consumers with Google AI Pro or Ultra subscriptions, as well as to developers and enterprise customers through Google's standard API channels.
What this means
Google's messaging here is unusually candid: a model can get more capable while quietly costing more, even without a price hike on paper. The distinction between per-token pricing and total task cost matters more as models make more tool calls and take more "turns" to complete agentic work — a trend that will likely spread across the industry as reasoning-heavy models become standard. The gated Fairwind Program and Cyber variant also signal that Google is treating offensive cyber and CBRN capabilities as categories requiring separate access controls rather than blanket public availability, a pattern likely to become more common as frontier models grow more capable at security-relevant tasks.
Related Articles
Google Rolls Out Gemini 3.8 Flash, Third Flash Update in Three Months
Google has released Gemini 3.8 Flash, the third Flash-tier update in three months, arriving just three weeks after Gemini 3.7 Flash. The model is live now in the Gemini app, Google Antigravity, and AI Studio with introductory pricing of $0.75/1M input and $3.75/1M output tokens.
Anthropic Releases Claude Fable 5.1 and Mythos 5.1, Cuts Cache Pricing 75% But Output Tokens Jump 70%
Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, claiming the top spot on Artificial Analysis's Intelligence Index at 66. Cache-read pricing dropped 75% to $0.25 per million tokens, but a 1.7x increase in output token usage pushes net per-task cost up 20%.
Google Adds Agent-Based Video Analysis to Gemini Flash, Cutting Token Usage by Up to 88 Percent
Google is rolling out agent-based video analysis for Gemini Flash models that dynamically searches footage instead of scanning frame by frame. Google claims the approach cuts token usage by up to 88 percent and costs by 66 percent while improving accuracy, with no added API fee.
Google Adds Agentic Video Understanding to Gemini, Cutting Token Use by Up to 88%
Google DeepMind has launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, letting models dynamically scan video segments instead of processing at a fixed frame rate. The company claims the feature cuts token consumption by up to 88%, reduces costs by up to 66%, and improves accuracy by up to 7%.
Comments
Loading...