Google DeepMind releases Nano Banana 2 Lite at $0.034 per 1K image with 4-second generation, opens Gemini Omni Flash API
Google DeepMind released Nano Banana 2 Lite (gemini-3.1-flash-lite-image), its fastest image generation model with 4-second text-to-image latency priced at $0.034 per 1K-resolution image. The company also opened developer access to Gemini Omni Flash (gemini-omni-flash-preview) for video generation and editing at $0.10 per second of output.
Google DeepMind releases Nano Banana 2 Lite at $0.034 per 1K image with 4-second generation, opens Gemini Omni Flash API at $0.10 per video second
Google DeepMind released Nano Banana 2 Lite (gemini-3.1-flash-lite-image), its fastest image generation model with 4-second text-to-image latency priced at $0.034 per 1K-resolution image. The company also opened developer access to Gemini Omni Flash (gemini-omni-flash-preview) for video generation and editing at $0.10 per second of output.
Nano Banana 2 Lite targets high-throughput developer workflows where speed and cost matter more than maximum quality. According to Google, it's positioned as the recommended replacement for the original Nano Banana (gemini-2.5-flash-image).
Nano Banana 2 Lite specifications
- Latency: 4 seconds for text-to-image generation
- Pricing: $0.034 per 1K-resolution image
- API name: gemini-3.1-flash-lite-image
- Availability: Google AI Studio, Gemini API, Gemini Enterprise Agent Platform
Google claims the model maintains "reliable prompt adherence, strong character consistency and legible in-image text rendering" despite optimizations for speed. The company published internal benchmarks comparing Elo quality scores, latency, and cost against competitor models, though independent verification is not yet available.
The Nano Banana family now includes four models:
- Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image): Speed-optimized
- Nano Banana 2 (Gemini 3.1 Flash Image): Balanced performance
- Nano Banana Pro (Gemini 3 Pro Image): Quality-optimized for professional use
- Nano Banana (Gemini 2.5 Flash Image): Legacy model, upgrade recommended
Gemini Omni Flash video capabilities
Gemini Omni Flash, first introduced at Google I/O, is now accessible via API in public preview. The model handles video generation and editing from text, image, and video inputs with conversational refinement.
Key specifications:
- Pricing: $0.10 per second of video output (matching Veo 3.1 Fast)
- API name: gemini-omni-flash-preview
- Current output length: 10 seconds (longer durations coming)
- Availability: Google AI Studio, Gemini API
Current limitations:
- Video references up to 3 seconds accepted but not correctly processed
- Audio reference uploads not yet supported
- Character consistency issues across scene changes
- Scene extension not available
The model uses Gemini's multimodal reasoning to synchronize text, graphics, and actions in generated video. Google positions it for conversational editing workflows where users iteratively refine outputs through natural language.
Combined workflow capabilities
Google published three demo applications showing Nano Banana 2 Lite generating initial images that Gemini Omni Flash then animates:
- Anywhere: Generates landmark backgrounds, animates into video clips
- Space Lift: Interior design visualization with cinematic walkthroughs
- Omni Product Studio: E-commerce product video generation
The Interactions API supports up to three sequential edits with maintained session context. Both models include SynthID watermarking for content provenance.
What this means
Google is aggressively pricing Nano Banana 2 Lite to compete in the high-volume image generation market, undercutting alternatives where speed matters more than quality. The $0.034 per 1K image pricing and 4-second latency target prototyping and draft-heavy workflows. However, without independent benchmarks, developers need to validate quality claims against their specific use cases.
The Gemini Omni Flash API opening is more significant. At $0.10 per second, it's the first major video generation model with conversational editing officially available via API, though the 10-second limit and processing bugs (especially around video references) suggest it's genuinely preview-stage. The combined image-to-video pipeline could enable new product categories in e-commerce, marketing, and content creation—if the quality and reliability hold up at scale.
Related Articles
Google Releases Gemini 3.7 Flash, Cuts Price in Half Versus 3.6 Flash
Google has released Gemini 3.7 Flash, just three weeks after Gemini 3.6 Flash, claiming substantial gains in coding, web development, and document reasoning. The model launches at an introductory price of $0.75 per 1M input tokens and $3.75 per 1M output tokens — half the cost of its predecessor.
xAI's Grok 4.6 Matches Claude and GPT-5.6 on Benchmarks, Costs 60% Less
xAI's Grok 4.6 ties OpenAI's GPT-5.6 Sol on the Artificial Analysis Intelligence Index with a score of 61, trailing only Anthropic's Claude Opus 5 and Claude Fable 5. Pricing remains at $2/$6 per million tokens, undercutting both competitors by more than 60 percent.
Alibaba Releases Qwen3.8 Open-Weight Models Under Apache 2.0, Including 27B Multimodal Model with 262K Native Context
Alibaba's Qwen team has released open weights for Qwen3.8, including a 27-billion-parameter multimodal dense model with 262,000 tokens of native context. The models ship under the Apache 2.0 license and are available on Hugging Face and ModelScope.
Alibaba Releases Qwen3.8-27B-FP8, a 27B Dense Vision-Language Model with 1M-Token Context
Alibaba's Qwen team has released FP8-quantized weights for Qwen3.8-27B, a 27-billion-parameter dense vision-language model with native 262,144-token context extensible to 1 million tokens. The model claims gains over its Qwen3.6 and Qwen3.7 predecessors on coding, agentic, and multimodal benchmarks.
Comments
Loading...