Google Ships Gemini 3.7 Flash, Cuts Price 50% Just Three Weeks After 3.6 Flash
Google released Gemini 3.7 Flash just three weeks after its predecessor, posting sharp gains on coding benchmarks while cutting launch pricing in half to $0.75 per million input tokens and $3.75 per million output tokens.
Google has released Gemini 3.7 Flash, only three weeks after shipping Gemini 3.6 Flash. The company positions the new model as its most capable workhorse for coding and AI agent tasks, while simultaneously cutting launch pricing by 50 percent compared to its predecessor.
Benchmark Gains
Google attributes the jump to "algorithmic improvements" rather than a larger model or new training run. The company reports the biggest gains in code quality:
- FrontierCode: 43.6 percent, up from 34.4 percent for Gemini 3.6 Flash
- DeepSWE: 65.3 percent, up from 49.0 percent
Google also claims improvements in web development, document comprehension, and business process automation, though it did not publish specific scores for those categories.
According to Google's own measurements, Gemini 3.7 Flash outperforms both Claude Sonnet 5 and GPT-5.6 Terra on these coding benchmarks. These figures come from Google's internal testing and have not been independently verified.
Pricing and Availability
Gemini 3.7 Flash is available now through the Gemini API, Google AI Studio, and Antigravity. Launch pricing is set at $0.75 per million input tokens and $3.75 per million output tokens — half the launch price of Gemini 3.6 Flash. The two models now share an identical price point.
Google says this pricing will hold through the end of the year. The models themselves may not last that long: Gemini 3.6 Flash survived just three weeks before being superseded.
Google has not disclosed the model's context window, parameter count, or training data cutoff date for Gemini 3.7 Flash.
What This Means
A three-week gap between Flash releases signals that Google is treating its lightweight model line as a rapid-iteration testbed rather than a stable product tier. For developers building on the Gemini API, this cadence means benchmark leadership claims have a short shelf life — Gemini 3.6 Flash's coding scores were already outdated before most teams finished evaluating it.
The 50 percent price cut is the more consequential detail. By matching 3.6 Flash's price floor while improving performance, Google is pushing harder on price-to-performance as the primary battleground in the sub-frontier model tier, where Anthropic's Claude Haiku line and OpenAI's smaller GPT-5.6 variants compete directly. Google's comparison against Claude Sonnet 5 and GPT-5.6 Terra — both larger, more expensive models — suggests an attempt to position a cheap Flash model as competitive with mid-tier frontier offerings on coding tasks specifically.
The caveat is that all comparative figures come from Google, measured on Google's chosen benchmarks. Independent verification, particularly on real-world coding agent tasks rather than benchmark suites like FrontierCode and DeepSWE, will determine whether the claimed gains hold up in production use.
Related Articles
Google Adds Real-Time Video Avatars to Gemini 3.8 Live for Enterprise Voice Agents
Google has added Live Avatar to its Gemini 3.8 Live dialogue models, pairing real-time video personas with voice AI for enterprise customer service and sales use cases. The feature is restricted to Gemini Enterprise customers via allowlisting and includes SynthID watermarking on all generated video.
OpenAI Cuts GPT-6 Sol and Luna Prices in Half, but Independent Benchmarks Show Flat Performance
OpenAI's GPT-6 Sol and Luna cut input/output token prices in half versus GPT-5.6, with Sol now at $2/$10 per million tokens and Luna at $0.10/$0.50. Independent testing from Artificial Analysis shows intelligence scores barely moved, with regressions on some knowledge-work benchmarks.
Anthropic Releases Claude Opus 5.5, Cuts Output Pricing to $20 per Million Tokens
Anthropic released Claude Opus 5.5 on Tuesday, cutting output token pricing to $20 per million tokens from $25 while improving coding and knowledge-work performance. The model arrives as Anthropic CEO Dario Amodei has pledged to slow capability advances to match safety work.
Vercel AI SDK Patch Adds Support for 'None' Reasoning Effort on GPT-6 Sol and Luna
Vercel released version 4.0.78 of @ai-sdk/openai, a patch that adds support for setting reasoningEffort to 'none' for GPT-6 Sol and Luna models. The update also introduces validation logic that warns on unsupported request-level effort updates and rejects invalid historical ones.
Comments
Loading...