Google Ships Gemini 3.7 Flash, Cuts Price 50% Just Three Weeks After 3.6 Flash
Google released Gemini 3.7 Flash just three weeks after its predecessor, posting sharp gains on coding benchmarks while cutting launch pricing in half to $0.75 per million input tokens and $3.75 per million output tokens.
Google has released Gemini 3.7 Flash, only three weeks after shipping Gemini 3.6 Flash. The company positions the new model as its most capable workhorse for coding and AI agent tasks, while simultaneously cutting launch pricing by 50 percent compared to its predecessor.
Benchmark Gains
Google attributes the jump to "algorithmic improvements" rather than a larger model or new training run. The company reports the biggest gains in code quality:
- FrontierCode: 43.6 percent, up from 34.4 percent for Gemini 3.6 Flash
- DeepSWE: 65.3 percent, up from 49.0 percent
Google also claims improvements in web development, document comprehension, and business process automation, though it did not publish specific scores for those categories.
According to Google's own measurements, Gemini 3.7 Flash outperforms both Claude Sonnet 5 and GPT-5.6 Terra on these coding benchmarks. These figures come from Google's internal testing and have not been independently verified.
Pricing and Availability
Gemini 3.7 Flash is available now through the Gemini API, Google AI Studio, and Antigravity. Launch pricing is set at $0.75 per million input tokens and $3.75 per million output tokens — half the launch price of Gemini 3.6 Flash. The two models now share an identical price point.
Google says this pricing will hold through the end of the year. The models themselves may not last that long: Gemini 3.6 Flash survived just three weeks before being superseded.
Google has not disclosed the model's context window, parameter count, or training data cutoff date for Gemini 3.7 Flash.
What This Means
A three-week gap between Flash releases signals that Google is treating its lightweight model line as a rapid-iteration testbed rather than a stable product tier. For developers building on the Gemini API, this cadence means benchmark leadership claims have a short shelf life — Gemini 3.6 Flash's coding scores were already outdated before most teams finished evaluating it.
The 50 percent price cut is the more consequential detail. By matching 3.6 Flash's price floor while improving performance, Google is pushing harder on price-to-performance as the primary battleground in the sub-frontier model tier, where Anthropic's Claude Haiku line and OpenAI's smaller GPT-5.6 variants compete directly. Google's comparison against Claude Sonnet 5 and GPT-5.6 Terra — both larger, more expensive models — suggests an attempt to position a cheap Flash model as competitive with mid-tier frontier offerings on coding tasks specifically.
The caveat is that all comparative figures come from Google, measured on Google's chosen benchmarks. Independent verification, particularly on real-world coding agent tasks rather than benchmark suites like FrontierCode and DeepSWE, will determine whether the claimed gains hold up in production use.
Related Articles
Meta Ships Muse Spark 1.2 Coding Model and Muse Code Agent, Undercuts Rivals with $0.20 Output Pricing
Meta released Muse Spark 1.2, a coding-focused upgrade to Spark 1.1, alongside Muse Code, its first dedicated terminal coding agent. The cheapest pricing tier drops output tokens to $0.20 per million, but requires users to share their data for training.
Google Assistant Shuts Down on Android and Wear OS Starting September 4, 2026
Google has set September 4, 2026 as the start date for removing Google Assistant access on Android phones, tablets, Wear OS watches, headphones, and Android Auto. Gemini becomes the sole assistant on these platforms, though Google Assistant remains active in cars with Google built-in.
OpenAI Launches 'Ultrafast' Mode, Claims 14x Speed Boost for GPT 5.6 Sol via Cerebras Partnership
OpenAI has introduced 'Ultrafast,' a preview mode that it claims accelerates GPT 5.6 Sol to 14 times standard speed, hitting up to 750 output tokens per second. The feature runs on OpenAI's partnership with chipmaker Cerebras and is currently limited to a small group of customers.
OpenAI Previews 'Ultrafast' Tier for GPT-5.6 Sol, Claims Up to 14x Speed Increase
OpenAI is testing an 'Ultrafast' service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing, generating up to 750 output tokens per second using Cerebras infrastructure. Access is currently limited to a waitlist of select customers.
Comments
Loading...