model release

Google Lists Gemini 3.8 Flash on OpenRouter With 1M-Token Context, September 2026 Release Date

TL;DR

Google's Gemini 3.8 Flash has surfaced on OpenRouter with a 1-million-token context window and discounted pricing of $0.75 per 1M input tokens and $3.75 per 1M output tokens. Google has not issued a separate public announcement, and the listed release date of September 2, 2026 is unusually far out, leaving key details unconfirmed.

2 min read
0

Gemini 3.8 Flash — Quick Specs

Context window1000K tokens
Input$0.75/1M tokens
Output$3.75/1M tokens

Gemini 3.8 Flash Appears on OpenRouter

Google's Gemini 3.8 Flash is now listed on OpenRouter, the model-routing marketplace, with a context window of 1 million tokens and a discounted price of $0.75 per 1M input tokens and $3.75 per 1M output tokens. OpenRouter's listing marks the price as "50% off," implying a standard rate of roughly $1.50 per 1M input tokens and $7.50 per 1M output tokens once the promotion ends.

As of publication, Google has not issued a corresponding blog post, developer changelog, or press release confirming the model independently of the OpenRouter listing. The information in this article is sourced entirely from that listing.

What's Known

According to the OpenRouter product page, Gemini 3.8 Flash is described as "Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." That description is a company/marketplace claim, not an independently verified benchmark result — no MMLU, HumanEval, or other standard benchmark scores were published alongside the listing.

Confirmed specifications from the listing:

  • Context window: 1,000,000 tokens
  • Input pricing: $0.75 per 1M tokens (listed as 50% off)
  • Output pricing: $3.75 per 1M tokens (listed as 50% off)
  • Listed release date: September 2, 2026
  • OpenRouter uptime (24-hour window): 100%

The September 2, 2026 release date is notable because it falls well in the future relative to typical model-release reporting conventions, and it has not been corroborated by any Google-owned channel. It's possible this is a placeholder date in OpenRouter's system, a scheduling artifact, or a genuine forward-dated listing tied to a staged rollout. Readers should treat the date as unconfirmed until Google publishes its own documentation.

No training cutoff date, parameter count, or modality breakdown (text-only vs. multimodal input/output) was disclosed in the source listing. Given that Gemini models have historically supported text, image, and in some cases audio and video input, Gemini 3.8 Flash is presumed multimodal, but this has not been explicitly confirmed for this specific version.

What This Means

This listing gives builders an early pricing and context-window signal for Gemini 3.8 Flash before any formal Google announcement — a pattern that has become common as model marketplaces like OpenRouter surface new checkpoints ahead of official vendor blog posts. The 1M-token context window keeps pace with Google's Gemini 1.5/2.x Flash line, and the sub-$1 input pricing keeps Flash positioned as the low-cost, high-throughput tier of Gemini's lineup rather than a frontier model competing on raw benchmark scores.

The unresolved release date is the story's biggest open question. Until Google confirms Gemini 3.8 Flash through its own channels — with a verifiable release date, benchmark results, and modality details — developers building on this model via OpenRouter should treat the listing as early access information rather than a finalized product announcement.

Related Articles

model release

Inception Launches Mercury 2.5 Preview, a Diffusion LLM Claiming 1,107 Tokens/Sec

Inception released Mercury 2.5 Preview, a diffusion-based language model that generates tokens in parallel rather than sequentially, claiming throughput of 1,107 tokens per second on standard GPUs. The model is available on OpenRouter with a 260K context window and an 80% launch discount through September 8, 2026.

model release

Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents

Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.

model release

Google Launches Gemini 3.5 Transcribe with 4.0% Word Error Rate Across 85 Languages

Google has released Gemini 3.5 Transcribe, a speech-to-text model that automatically detects 85 languages, removes filler words, and corrects misspoken phrases. The company claims a 4.0 percent word error rate for streaming audio and 70 percent lower latency than its predecessor, Chirp 3.

model release

Google DeepMind Ships Gemini 3.8 Flash and a Cybersecurity Variant, Third Flash Release in Six Weeks

Google DeepMind released Gemini 3.8 Flash and a specialized cybersecurity variant, Gemini 3.8 Flash Cyber, its third Flash-tier launch in six weeks. Pricing stays at $0.75 per million input tokens and $3.75 per million output tokens, matching the prior 3.7 Flash release.

Comments

Loading...