image-generation
43 articles tagged with image-generation
Microsoft Releases Mage-Flow: Compact 4B Image Generation and Editing Models Matching Systems 5-8x Larger
Microsoft has released Mage-Flow, a family of 4B-parameter image generation and editing models built on a shared tokenizer-transformer stack. According to Microsoft, the Turbo variants match or beat open-source systems with 5-8x more parameters while running in 4 diffusion steps.
Black Forest Labs Reports 10x Fewer Safety Vulnerabilities Than Competitors in FLUX.2 Model Family
Black Forest Labs reports its FLUX.2 image generation models demonstrate more than 10 times fewer vulnerabilities for synthetic non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM) compared to other leading open-weight models. The company claims targeted post-training mitigations reduced vulnerabilities by 77-98% before release, according to third-party red-teaming conducted by Cinder.
Black Forest Labs releases FLUX.2: 32B open-weight image model with 4MP editing and 10-image multi-reference support
Black Forest Labs has released FLUX.2, a family of image generation models including a 32B parameter open-weight variant. The models support editing at up to 4 megapixel resolution and can reference up to 10 images simultaneously for character and style consistency.
Black Forest Labs' FLUX.1 Kontext [Pro] Now Available in Adobe Photoshop Beta
Black Forest Labs' FLUX.1 Kontext [Pro] is now available inside Adobe Photoshop's Generative Fill feature, starting September 25. The company claims the model is 3x faster than competing generative fill models and will be free to use during the beta period.
Black Forest Labs Releases FLUX.1 Krea [dev], Open-Weights Image Model Trained to Avoid 'Oversaturated AI Look'
Black Forest Labs released FLUX.1 Krea [dev], an open-weights text-to-image model developed with Krea AI. The model was specifically trained to generate realistic images that avoid the oversaturated textures common in AI-generated imagery, achieving performance on par with FLUX1.1 [pro] in human preference tests.
Meta Auto-Opts All Public Instagram Accounts Into AI Image Remixing Without Notification
Meta launched its Muse Image AI generator for Instagram, automatically opting in all public accounts to allow others to tag and remix their photos without notification. The feature, which lets anyone generate AI images based on a user's likeness by tagging their account, can be disabled through Instagram's Sharing and reuse settings.
Meta Enables AI Image Generation Using Any Public Instagram Profile Photo by Default
Meta launched Muse Image, an AI image generator integrated into Instagram that uses public profile photos and posts as prompts by default. Users can only disable this by manually switching off two toggles buried in the app's privacy settings.
Google releases Gemini 3.1 Flash Lite Image, its fastest and cheapest image generation model
Google has released Gemini 3.1 Flash Lite Image, also called Nano Banana 2 Lite, which the company describes as its fastest and cheapest image generation model. The model is available through Google's AI Studio and Gemini API with the identifier gemini-3.1-flash-lite-image.
Google launches Nano Banana 2 Lite image model at 4 seconds per image, $0.04 per 1,000 generations
Google released Nano Banana 2 Lite, an image generation model that produces images in four seconds at under four cents per thousand images. The model prioritizes speed and cost over quality, targeting developers building high-volume image pipelines.
Google launches Gemini 3.1 Flash Lite Image with 4-second generation time, $0.25 per 1M input tokens
Google has released Gemini 3.1 Flash Lite Image, a text-to-image model that generates 1K resolution images in approximately 4 seconds — 2.7× faster than Gemini 3.1 Flash Image. The model is priced at $0.25 per 1M input tokens and $1.50 per 1M output tokens, with a 66K context window and knowledge cutoff of January 2025.
Google DeepMind releases Nano Banana 2 Lite at $0.034 per 1K image with 4-second generation, opens Gemini Omni Flash API
Google DeepMind released Nano Banana 2 Lite (gemini-3.1-flash-lite-image), its fastest image generation model with 4-second text-to-image latency priced at $0.034 per 1K-resolution image. The company also opened developer access to Gemini Omni Flash (gemini-omni-flash-preview) for video generation and editing at $0.10 per second of output.
Proton launches Lumo 2.0 with multimodal capabilities, scores 240% higher on AI benchmarks
Proton has released Lumo 2.0, adding image recognition and generation, encrypted memory features, and enhanced web search to its privacy-focused AI assistant. The company claims Lumo 2.0 Max scored 240% higher than version 1.4 on the Artificial Analysis Intelligence Index, while maintaining zero-access encryption and no conversation logging.
Google brings personalized image generation to all US Gemini users, expanding from paid-only feature
Google is expanding personalized image generation in the Gemini app to all eligible US users, removing the previous restriction to AI Pro and Ultra subscribers. The feature allows Gemini to access user data across Google services like Gmail and Photos when generating images.
Google makes Gemini's personalized image generation free for all U.S. users
Google removed the paywall for Gemini's personalized image generation feature, making it free for all eligible U.S. users starting today. The Nano Banana-powered feature was previously limited to Plus, Pro, and Ultra subscribers.
Unconventional AI releases Un-0 image model on simulated oscillator chip claiming 1000x power reduction
Unconventional AI released Un-0, an image generation model that runs on a software simulation of oscillator-based hardware. Founder Naveen Rao claims the architecture could reduce AI power consumption by 1000x compared to conventional chips, though no physical hardware exists yet.
Unconventional AI releases Un0 image model on oscillator-based architecture, claims 1,000x power reduction potential
Unconventional AI, led by former Databricks AI chief Naveen Rao, has released Un0, an image generation model built on a simulated oscillator-based architecture. The company claims this approach could reduce inference power consumption by up to 1,000x compared to conventional computing, though the technology currently runs only in software simulation.
Mistral Launches Free Chat Interface With Web Search, Canvas Editor, and Image Generation
Mistral AI has released major updates to le Chat, its free AI assistant, adding web search with citations, a canvas interface for collaborative editing, and image generation powered by Black Forest Labs' Flux Pro. The updates are powered by the new Pixtral Large multimodal model and are available free in beta.
Google releases Gemini 3.1 Flash Image, claims Pro-level quality at $0.50 per 1M tokens
Google has released Gemini 3.1 Flash Image, internally codenamed "Nano Banana 2," an image generation and editing model with a 131K context window. The model is priced at $0.50 per 1M input tokens and $3 per 1M output tokens.
Apple Intelligence adds cross-app context awareness, one-tap password fixes, and AI-generated Safari extensions
Apple announced a suite of Apple Intelligence updates across iOS apps, including cross-app context awareness that surfaces email details during phone calls, one-tap AI-powered password updates, and natural language Shortcuts creation. Safari gains AI-generated extensions, while Photos gets spatial reframing and improved object removal.
Apple adds cloud-powered AI image editing to iOS 27 Photos app with Clean Up upgrade, Extend and Reframe tools
Apple announced three AI-powered photo editing tools for iOS 27: an upgraded Clean Up feature for object removal, Extend for generating content around image borders, and Reframe for changing photo angles. The features use a combination of on-device and cloud AI models, marking a shift from Apple's previous on-device-only approach.
Microsoft releases MAI-Thinking-1, its first reasoning model with 35B parameters
Microsoft released seven AI models at Build 2026, headlined by MAI-Thinking-1, its first reasoning model with 35 billion parameters. The company claims the model matches Anthropic's Claude Opus 4.6 on SWE Bench Pro coding benchmarks and beats Sonnet 4.61 in blind tests.
Apple to upgrade on-device image models in iOS 27, add third-party AI image generation support
Apple plans to significantly improve the visual quality of its on-device image generation models for Genmoji and Image Playground in iOS 27, according to Bloomberg's Mark Gurman. The update will also add support for third-party AI image generation models beyond OpenAI's ChatGPT.
OpenAI adopts C2PA metadata standard and Google's SynthID watermarking for AI image detection
OpenAI is joining the C2PA open standard and embedding Google DeepMind's invisible SynthID watermark in all AI-generated images from its models. The company is launching a public verification tool that checks for both C2PA metadata and SynthID watermarks, though detection only works for images created by OpenAI's own products.
Image AI models drive 6.5x more app downloads than text model updates, Appfigures data shows
Image model releases are generating 6.5 times more mobile app downloads than traditional text model updates, according to Appfigures. Google's Gemini added 22 million downloads in 28 days following its image model release, while ChatGPT added 12 million after GPT-4o image capabilities launched.
ChatGPT Images 2.0 Adds UI Design Analysis and Mockup Generation Capabilities
OpenAI's ChatGPT Images 2.0 has added UI design analysis capabilities, allowing it to review interface designs, flag specific issues, and generate redesigned mockups. The feature is available to ChatGPT Plus subscribers at $20/month and represents an expansion beyond pure image generation into design review.
ChatGPT Images 2.0 scores 97% in head-to-head image generation benchmark against Google's Gemini Nano Banana at 85%
OpenAI's ChatGPT Images 2.0 scored 97% versus Google's Gemini Nano Banana at 85% in a nine-test image generation benchmark conducted by ZDNET. The tests measured capabilities including image restoration, text rendering, and prompt adherence, with Nano Banana losing points primarily for fabricating details and text errors.
OpenAI releases ChatGPT Images 2.0 with 3840x2160 resolution at $30 per 1M output tokens
OpenAI released ChatGPT Images 2.0, pricing output tokens at $30 per million with maximum resolution of 3840x2160 pixels. CEO Sam Altman claims the improvement from gpt-image-1 to gpt-image-2 equals the jump from GPT-3 to GPT-5.
OpenAI launches ChatGPT Images 2 with 2K resolution and two-mode generation
OpenAI has released ChatGPT Images 2, an upgraded image generation model that produces images up to 2K resolution in multiple aspect ratios. The model ships with two versions—Instant and Thinking—and can research current web information before generating images.
OpenAI announces gpt-image-2 model with improved text rendering and UI generation
OpenAI is set to announce gpt-image-2, its next-generation image generation model, on April 21, 2026 at 12pm PT. The company's teaser demonstrates improved capabilities in rendering text and generating realistic user interfaces from text prompts.
Google adds Nano Banana image generation to Gemini Personal Intelligence, using Gmail and Photos data
Google has integrated its Nano Banana image generation system with Gemini's Personal Intelligence feature, enabling the AI to create images informed by user data from Gmail, Photos, Calendar, Drive, and other Google apps. The feature rolls out to Plus, Pro, and Ultra subscribers in the US first, with Europe excluded from the initial launch.
Google's Gemini now generates personalized images using your Google Photos library
Google's Gemini can now generate personalized images by pulling data from users' Google Photos libraries through its Personal Intelligence feature. The integration uses Google Photos labels to identify people and objects, then generates images via the Nano Banana 2 model that reflect users' tastes and lifestyle.
Baidu releases ERNIE-Image-Turbo, a distilled text-to-image model generating in 8 inference steps
Baidu has released ERNIE-Image-Turbo, a distilled text-to-image diffusion transformer that generates images in 8 inference steps. The model runs on consumer GPUs with 24GB VRAM and supports resolutions up to 1376×768, with claimed strengths in text rendering and structured generation tasks.
Stability AI launches Brand Studio for enterprise image generation with brand-specific models
Stability AI has launched Brand Studio, a commercial platform designed for creative teams to generate AI images aligned with their brand identity. The platform includes Brand Central for training custom models, Producer Mode for automated visual workflows, and Curated Model Routing that selects optimal models for specific tasks.
Stability AI and NVIDIA launch Stable Diffusion 3.5 NIM for faster image generation
Stability AI and NVIDIA have launched Stable Diffusion 3.5 NIM, a microservice designed to accelerate image generation performance and simplify enterprise deployment. The collaboration packages Stable Diffusion 3.5 as an NVIDIA NIM (NVIDIA Inference Microservice) for optimized inference.
Stable Diffusion 3.5 TensorRT optimization delivers 2x faster generation, 40% less VRAM on RTX GPUs
Stability AI has released TensorRT-optimized versions of the Stable Diffusion 3.5 model family in collaboration with NVIDIA. The optimization uses FP8 quantization to achieve 2x faster generation speed and 40% lower VRAM requirements on supported RTX GPUs.
Stable Diffusion optimized for AMD Radeon GPUs and Ryzen AI APUs
Stability AI has released ONNX-optimized versions of Stable Diffusion engineered to run faster and more efficiently on AMD Radeon GPUs and Ryzen AI APUs. The collaboration with AMD targets broader hardware compatibility for the image generation model.
Stable Diffusion 3.5 Large launches on Microsoft Azure AI Foundry
Stability AI's Stable Diffusion 3.5 Large model is now available through Microsoft Azure AI Foundry, giving businesses integrated access to professional-grade image generation within Azure's ecosystem. The deployment expands SD3.5 Large's availability across major cloud platforms.
Adobe Firefly now learns custom visual styles from user-uploaded images
Adobe is rolling out custom models for Firefly, allowing creators to train the generative model on 10-30 of their own images to generate new content matching their specific visual style. The feature costs 500 credits per training session and supports three methods: photography style, illustration style, and character consistency.
Microsoft's superintelligence team releases MAI-Image-2, ranks third in text-to-image generation
Microsoft's superintelligence team, led by Mustafa Suleyman, has released MAI-Image-2, a text-to-image generator that currently ranks third on the Arena.ai leaderboard for text-to-image models, behind OpenAI's GPT-Image-1.5 and Google's Nano Banana 2. The model is now available for testing in the MAI Playground and will roll out to Copilot and Bing Image Creator, with API access opening to all developers through Microsoft Foundry.
Midjourney V8 achieves 5x faster generation but premium features cost 4x more
Midjourney has released an early version of V8 for community testing, achieving roughly 5x faster image generation and introducing native 2K resolution via --hd mode. However, premium features including --hd, --q 4, style references, and mood boards cost four times as much as standard generation, with Relax mode unavailable at launch.
Google DeepMind releases Nano Banana 2 image model with Pro-level capabilities at faster speeds
Google DeepMind has released Nano Banana 2, an image generation model that combines advanced world knowledge and subject consistency with faster inference speeds comparable to its Flash offering. The model is positioned as production-ready with capabilities previously associated with Pro-tier performance.
Google relaunches Flow AI studio with free image generation and video editing
Google has relaunched its Flow AI creative studio as a unified platform for image and video creation. The updated tool includes free image generation capabilities and new editing features designed to streamline creative workflows.
Segmind releases SegMoE, a mixture-of-experts diffusion model for faster image generation
Segmind has released SegMoE, a mixture-of-experts (MoE) diffusion model designed to accelerate image generation while reducing computational overhead. The model applies MoE techniques traditionally used in large language models to the diffusion model architecture, enabling selective expert activation during inference.