Google caps single-prompt quota for Gemini 3.1 Pro, makes Flash-Lite free after usage limit complaints
Google has modified Gemini's compute-based usage limits introduced at I/O 2026 after users reported depleting quotas too quickly. The company is now capping how much quota a single Gemini 3.1 Pro prompt can consume and making all 3.1 Flash-Lite prompts free.
Google caps single-prompt quota for Gemini 3.1 Pro, makes Flash-Lite free after usage limit complaints
Google has adjusted Gemini's new compute-based usage limits one week after their introduction at I/O 2026, responding to user complaints about hitting quotas too quickly.
Key changes
The company is now capping the amount of quota a single prompt can use when accessing Gemini 3.1 Pro, according to Gemini lead Josh Woodward. This addresses issues where complex prompts with large files were rapidly depleting user limits.
Google has made all Gemini 3.1 Flash-Lite prompts free, with these requests no longer counting against user quotas.
The company also doubled the number of Omni generations available to Google AI Ultra subscribers after fixing a bug where "just one or two Omni videos" would drain quotas for some users.
How the compute-based system works
Google introduced compute-based usage limits at I/O 2026 to replace previous approaches. The system accounts for prompt complexity, tools used, and chat length, with a 5-hour refresh period until weekly limits are met. According to Google, "a simple text prompt uses far less compute than a complex video or coding prompt."
Failed requests now do not count against limits. Google stated: "If a request fails, you won't be charged. Our system mistakes are on us, not you. Your quota is used only for successful completions."
Additional improvements
Google plans to provide more detailed usage breakdowns and notifications for compute-intensive tasks like Deep Research. The current gemini.google.com/usage dashboard provides only a high-level overview.
The company will introduce pay-as-you-go top-up AI credits for users who need additional quota.
Google also confirmed that when users select a specific model, the system remembers that choice across future sessions unless manually changed or a quota cap triggers automatic fallback to a lighter model.
What this means
The rapid adjustments reveal that Google's initial compute-based limit implementation was too restrictive for real-world usage patterns. By capping per-prompt consumption and making the lighter Flash-Lite tier free, Google is attempting to balance resource management with user experience. The move to exempt Flash-Lite from quotas suggests Google wants to retain casual users while reserving compute limits primarily for its more powerful Pro models. The bug affecting Omni video generation indicates the new system launched with significant technical issues that required immediate correction.
Related Articles
Vercel AI SDK Patch Adds Support for Unreleased 'gemini-3.7-flash' Model ID
Vercel shipped a patch release of @ai-sdk/google-vertex (v4.0.182) that adds support for a model identifier called 'gemini-3.7-flash.' Google has not publicly announced this model, and no official details on pricing, context window, or benchmarks exist yet.
Cline Desktop v0.0.19 Fixes Memory Leak That Ballooned Process to Tens of Gigabytes
Cline Desktop v0.0.19 fixes a memory leak where session status updates carried full conversation transcripts to every connected client, ballooning process memory to tens of gigabytes on long tasks. The release also adds seven new model providers and changes default models for several existing ones.
Ollama v0.33.0 Fixes KV Cache Bug That Forced Full Reprocessing of 46K-Token Prompts
Ollama's v0.33.0 pre-release adds direct model management inside Claude Desktop's menu bar and fixes a caching bug that could force reprocessing of tens of thousands of tokens. The release also disables a Claude Code system message that was breaking KV cache hits on every request.
Cline v4.1.11 Adds Inline Image Generation, Fixes VS Code 1.134 Compatibility Break
Cline's v4.1.11 update adds inline image generation for supported models and fixes a code actions bug that broke on VS Code 1.134. The release also expands the model catalog with eight new providers including AMD, Arcee, and RunInfra.
Comments
Loading...