model release

Google Releases Gemini 3.7 Flash With 1M-Token Context and Multimodal Input

TL;DR

Google has released Gemini 3.7 Flash, a multimodal model built for agentic workflows, coding, and multi-step reasoning. It offers a 1,049K token context window and is priced at $0.38 per million input tokens and $1.88 per million output tokens, available now via OpenRouter.

2 min read
0

Gemini 3.7 Flash — Quick Specs

Context window1049K tokens
Input$1.5/1M tokens
Output$7.5/1M tokens

Google has released Gemini 3.7 Flash, a multimodal AI model positioned for fast agentic workflows, coding tasks, and complex multi-step reasoning. The model is now accessible through OpenRouter's API under the identifier google/gemini-3.7-flash.

Specifications

Gemini 3.7 Flash ships with a context window of 1,049,000 tokens, placing it among the largest-context models currently available. According to Google, the model is designed for tasks requiring responsive performance and reliable execution across multiple sequential steps — a profile aimed squarely at agentic and coding use cases rather than single-shot chat.

The model accepts text, image, file, audio, and video inputs and produces text outputs, making it a fully multimodal system rather than a text-only release.

Pricing

OpenRouter lists pricing for Gemini 3.7 Flash at $0.38 per million input tokens and $1.88 per million output tokens. That output price is roughly five times the input price, a common pattern for models tuned to reward concise, high-value generations over lengthy ones.

What's known and what isn't

Google's own description frames Gemini 3.7 Flash around agentic workflows, coding, and multi-step reasoning, but no independent benchmark scores were included in the release materials reviewed for this article. No parameter count, training data cutoff date, or detailed architecture information has been disclosed. Claims about the model's suitability for agentic tasks and coding come directly from Google's positioning and have not been independently verified with benchmark data at time of writing.

What this means

A 1,049K-token context window paired with sub-$2 output pricing signals Google is continuing to push the Flash line as a high-volume, cost-efficient option for developers building agents and coding tools that need to process large amounts of context — think entire codebases, long documents, or extended video/audio inputs — without paying frontier-model prices. The five-to-one output-to-input pricing ratio suggests Google expects most usage to be input-heavy (large context, short responses), which aligns with retrieval-augmented and agentic patterns where a model reads a lot but writes comparatively little.

The absence of published benchmark numbers makes it difficult to assess how Gemini 3.7 Flash compares to competing fast-tier models on coding or reasoning tasks. Developers evaluating it for production agentic pipelines will need to run their own tests until third-party benchmarks emerge. Availability via OpenRouter at launch also means the model is immediately accessible to a broad developer base without requiring direct Google Cloud onboarding.

Related Articles

model release

Meta Releases Muse Glimmer 30B, an Open-Weight Agentic Model for Consumer Hardware

Meta Superintelligence Labs has released Muse Glimmer 30B, a dense open-weight model distilled from its larger Muse Spark system and tuned for agentic workflows on consumer hardware. The model supports 131K context, image understanding, and over 100 languages at $0.30/$1.10 per 1M input/output tokens.

model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

model release

Google Launches Gemini 3.8 Flash TTS: Voice Cloning and Text-Described Voices for $9-18 per Million Audio Tokens

Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, two speech generation models that let users design voices from text descriptions or clone a voice from a 30-second sample. Both support over 100 languages and roll out now through the Gemini API and Google AI Studio.

model release

Anonymous Stealth Model "Space Bunny Alpha" Debuts on OpenRouter With 1M-Token Context, Free During Preview

A previously unknown AI provider has released Space Bunny Alpha, a stealth model on OpenRouter offering a 1M-token context window, adjustable reasoning effort, and multimodal input support. The model is free during its preview period, though its developer remains unnamed.

Comments

Loading...