model release

Google Releases Gemini 3.5 Flash with 1M Token Context and Configurable Thinking Modes at $1.50/$9 Per Million Tokens

TL;DR

Google has released Gemini 3.5 Flash, a multimodal model with a 1 million token context window priced at $1.50 per million input tokens and $9 per million output tokens. The model supports text, image, video, audio, and PDF inputs with configurable thinking effort levels from minimal to high.

2 min read
0

Gemini 3.5 Flash — Quick Specs

Context window1049K tokens
Input$1.5/1M tokens
Output$9/1M tokens

Google Releases Gemini 3.5 Flash with 1M Context and Thinking Modes

Google has released Gemini 3.5 Flash, a multimodal model priced at $1.50 per million input tokens and $9 per million output tokens. The model features a 1 million token context window and supports text, image, video, audio, and PDF inputs.

Key Specifications

Gemini 3.5 Flash is positioned as a high-efficiency model delivering what Google describes as "near-Pro level coding and reasoning at Flash-tier cost and speed." The model defaults to medium thinking effort for standard responses but supports four configurable thinking levels: minimal, low, medium, and high.

The thinking mode configuration allows developers to make explicit cost-performance trade-offs based on task complexity. This feature is designed for parallel agentic execution loops where different subtasks may require different computational resources.

Technical Capabilities

According to Google, the model is "highly optimized for coding proficiency" and multimodal processing. The 1 million token context window positions it for handling large codebases, extensive documentation, and long-form content analysis.

The multimodal capabilities extend across five input types: text, static images, video, audio, and PDF documents. This broad input support makes the model applicable to document processing, multimedia analysis, and complex reasoning tasks that span multiple data formats.

Pricing and Availability

At $1.50 per million input tokens and $9 per million output tokens, Gemini 3.5 Flash is priced competitively in the Flash model tier. The model is available through OpenRouter with routing to multiple providers for reliability and uptime optimization.

The release date is listed as May 19, 2026 in the source material, though this appears to be a future date and may represent a placeholder or projected availability timeline.

What This Means

Gemini 3.5 Flash's configurable thinking modes represent a shift toward explicit computational trade-offs in model inference. Rather than offering a single performance point, developers can adjust reasoning depth based on task requirements—a feature particularly relevant for agentic workflows where some operations need deep reasoning while others prioritize speed.

The 1M context window combined with multimodal support and competitive pricing positions this model for code analysis, document processing, and complex multi-step reasoning tasks. The thinking mode feature may influence how other providers structure their model offerings, particularly for use cases requiring variable computational intensity across different parts of a workflow.

Related Articles

model release

Google unveils Gemini 4 Argon at $2/$10 per 1M tokens, but access is limited to select users

Google unveiled Gemini 4 Argon on Wednesday with introductory pricing of $2 per 1M input tokens and $10 per 1M output tokens, matching OpenAI's discounted GPT-6.1 Sol. Access is restricted to select cybersecurity defenders and enterprise cloud customers, and Google says published rates will double later.

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

model release

Google Releases Gemini 4 Argon, Positions It as Most Powerful Model Yet With Cybersecurity Focus

Google has released Gemini 4 Argon, a new model the company calls its most powerful yet, with a specific focus on defensive cybersecurity work. The model is currently limited to select partners through Google's Fairwind Program, with no public pricing or context window details disclosed.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

Comments

Loading...