Google DeepMind's Gemini 3.1 Flash-Lite generates websites in real time, 2.5x faster than predecessor
Google DeepMind released Gemini 3.1 Flash-Lite, a model that generates functional websites in real time through a new pseudo-browser demo. The model achieves first response token 2.5 times faster than Gemini 2.5 Flash and outputs over 360 tokens per second, though output pricing has tripled from $0.40 to $1.50 per million tokens.
Gemini 3.1 Flash Lite — Quick Specs
Google DeepMind's Gemini 3.1 Flash-Lite generates websites in real time, 2.5x faster than predecessor
Google DeepMind released Gemini 3.1 Flash-Lite with substantially improved inference speed. The model achieves first response token generation 2.5 times faster than Gemini 2.5 Flash and sustains output of over 360 tokens per second, according to Google.
The company demonstrated the capability through a new pseudo-browser interface: users input a text prompt describing a desired webpage, and the model renders HTML and CSS in real time as it generates. A live demo is available free in Google AI Studio.
Performance trade-offs
The speed gains come with significant cost increases. Output pricing has more than tripled to $1.50 per million tokens, up from $0.40 per million tokens on the prior Flash version. Input pricing was not disclosed in available sources.
The model's website generation output shows consistency issues. Generated pages begin rendering correctly but content "quickly drifts into nonsense," according to assessments of the demo. Google suggests tight guardrails could enable practical use cases such as rapid UI mockup creation for design visualization.
Competitive positioning
According to Artificial Analysis benchmarking, Gemini 3.1 Flash-Lite outperforms larger models including Claude Opus 4.6 on certain multimodal tasks, though comprehensive benchmark scores were not disclosed.
The model became available in Google AI Studio and Vertex AI starting in early March 2026.
What this means
Gemini 3.1 Flash-Lite prioritizes inference speed over output cost—a deliberate trade-off that positions the model for latency-sensitive applications where user-facing response time matters more than token expenses. The website generation capability remains a novelty demonstration rather than production-ready tooling, but the speed metrics signal Google's focus on competing in the low-latency inference market where models like Claude Sonnet and smaller specialized models have gained traction.
Related Articles
Google DeepMind Launches Gemini 3.8 Live, Claims #1 Spot on Speech-to-Speech Benchmark
Google DeepMind has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-dialogue models that reason and execute background tasks without interrupting conversation. Google claims the Extended Thinking model ranks #1 on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6.
Google DeepMind's Dream-RSI Cuts AI Search Costs by Replaying Past Attempts Instead of Repeating Them
Google and DeepMind researchers introduced Dream-RSI, a method that lets AI agents test new search strategies by replaying recorded past attempts instead of running costly new computations. Tested on Gemini 3.1 Pro and Gemini 3.7 Flash across eight tasks, it matched or beat baselines while using far fewer attempts.
Google Launches Gemini 3.8 Live, Undercutting OpenAI's GPT-Live-1 on Price by Up to 70%
Google DeepMind released Gemini 3.8 Live and a reasoning-enhanced Extended Thinking variant for voice agents, pricing audio input at $0.005/minute versus OpenAI's $0.05/minute for GPT-Live-1. The Extended Thinking model tops the Artificial Analysis Speech-to-Speech Leaderboard with 82.6 percent.
OpenRouter Adds DeepSeek Flash Latest Alias With 1M-Token Context Window
OpenRouter has launched deepseek-flash-latest, a persistent endpoint that always points to the current DeepSeek Flash model. It offers a 1,049K token context window, text-and-image input, and pricing of $0.15 per 1M input tokens and $0.60 per 1M output tokens.
Comments
Loading...