model release

Google releases Gemini 3.5 Flash with 4x faster output and agentic capabilities, 3.5 Pro coming June

TL;DR

Google released Gemini 3.5 Flash today with 4x faster output token generation than competing frontier models while surpassing Gemini 3.1 Pro on coding, agentic, and multimodal benchmarks. The company announced Gemini 3.5 Pro will launch next month and introduced Gemini Omni, a new multimodal series that outputs video.

3 min read
0

Gemini 3.5 Flash — Quick Specs

Context window1049K tokens
Input$1.5/1M tokens
Output$9/1M tokens

Google releases Gemini 3.5 Flash with 4x faster output and agentic capabilities, 3.5 Pro coming June

Google released Gemini 3.5 Flash today at I/O 2026, claiming the model delivers 4x faster output tokens per second compared to other frontier models while maintaining Flash series pricing. According to Google, the model surpasses Gemini 3.1 Pro in coding, agentic, and multimodal benchmarks.

The model is available immediately in the Gemini app, Google Search, Antigravity 2.0, and via the Gemini API. Gemini 3.5 Pro is currently in testing and will launch next month.

New Gemini Omni multimodal series

Google introduced Gemini Omni, a new model family combining reasoning with content creation. Gemini Omni Flash accepts image, audio, video, and text inputs and outputs video "grounded in real-world knowledge" that can be edited. The model is rolling out to AI Plus, Pro, and Ultra subscribers in the Gemini app, Google Flow, and YouTube Shorts.

The company also launched Google Flow and Flow Music as standalone mobile apps, with Flow available on Android in beta and Flow Music launching on iOS first.

Gemini Spark agentic assistant

Google announced Gemini Spark, described as "your personal agent" that performs tasks autonomously. The system integrates with Gmail, Docs, and other Google Workspace apps, with third-party tool support via MCP (Model Context Protocol) coming this summer.

Gemini Spark will be available next week exclusively to Google AI Ultra subscribers in the US. The service transforms Gemini "from an assistant that can answer your questions into an active partner that does real work on your behalf," according to Google.

Pricing changes for Google AI plans

Google restructured its subscription tiers. AI Ultra now starts at $100 per month (previously $250), offering 5x higher usage limits than AI Pro. A new $200 tier replaces the previous $250 plan with identical capabilities.

The company is shifting from daily prompt limits to a "compute-used" model that accounts for prompt complexity, features used, and chat length. Limits refresh every five hours until reaching a weekly cap.

Search and productivity updates

AI Mode in Google Search now runs on Gemini 3.5 Flash. New "information agents" will monitor the web 24/7 for topics users specify, available to AI Pro and Ultra subscribers this summer. Google Search will also gain the ability to build custom dashboards and trackers for ongoing tasks.

Gmail Live, a conversational email search feature, rolls out to AI Pro and Ultra subscribers in the US this summer on Android and iOS. Docs Live for conversational document creation and editing launches simultaneously for the same subscriber tiers.

What this means

Google's 4x speed claim for Gemini 3.5 Flash positions it directly against Anthropic's Claude 3.5 Sonnet and OpenAI's GPT-4o in the fast inference tier. The introduction of Gemini Omni's video output capabilities represents a significant multimodal expansion, though real-world quality benchmarks remain to be seen. The restructured pricing and compute-based limits suggest Google is attempting to balance access with infrastructure costs as model capabilities increase. The MCP integration for Gemini Spark indicates Google is adopting industry standards for agent interoperability rather than building a proprietary ecosystem.

Related Articles

model release

OpenAI Ships GPT-6.1 Sol at DevDay 2026, Claims Near-Astra Performance at One-Fifth the Price

OpenAI's DevDay 2026 keynote introduced GPT-6.1 Sol, a mid-tier model priced at $2/$10 per million tokens that OpenAI claims delivers 'near-Astra intelligence' at a fraction of the cost. Independent benchmarks from Artificial Analysis and third-party testers show it trailing flagship Astra by roughly one point on the Intelligence Index while beating Opus 5.5 on cost-adjusted coding tasks.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

model release

Google unveils Gemini 4 Argon at $2/$10 per 1M tokens, but access is limited to select users

Google unveiled Gemini 4 Argon on Wednesday with introductory pricing of $2 per 1M input tokens and $10 per 1M output tokens, matching OpenAI's discounted GPT-6.1 Sol. Access is restricted to select cybersecurity defenders and enterprise cloud customers, and Google says published rates will double later.

model release

Ideogram 4.5 launches with native 2K output and four tiers from $0.008 to $0.22 per image

Ideogram has released Ideogram 4.5, an image model it claims edits only the area a user specifies and leaves the rest untouched. It offers four quality tiers from 0.8 to 22 cents per image, all at native 2K resolution, via the Ideogram platform and API. An open-weight release is promised but not yet dated.

Comments

Loading...