multimodal AI
23 articles tagged with multimodal AI
Tencent Releases WeMM-Embedding-9B, a Multimodal Embedding Model Built on Qwen3.5
Tencent has released WeMM-Embedding-9B, a 9-billion-parameter multimodal embedding model built on Qwen3.5 that produces 4,096-dimensional embeddings from text, images, video, and visual documents. The model reports state-of-the-art results on the MMEB-v2 and MMEB-v3 benchmarks and is released under Apache 2.0.
Google Launches Gemini 3.5 Transcribe with 4.0% Word Error Rate Across 85 Languages
Google has released Gemini 3.5 Transcribe, a speech-to-text model that automatically detects 85 languages, removes filler words, and corrects misspoken phrases. The company claims a 4.0 percent word error rate for streaming audio and 70 percent lower latency than its predecessor, Chirp 3.
Alibaba Releases Qwen3.8 Open-Weight Models Under Apache 2.0, Including 27B Multimodal Model with 262K Native Context
Alibaba's Qwen team has released open weights for Qwen3.8, including a 27-billion-parameter multimodal dense model with 262,000 tokens of native context. The models ship under the Apache 2.0 license and are available on Hugging Face and ModelScope.
Google Releases Gemini 3.7 Flash With 1M-Token Context and Multimodal Input
Google has released Gemini 3.7 Flash, a multimodal model built for agentic workflows, coding, and multi-step reasoning. It offers a 1,049K token context window and is priced at $0.38 per million input tokens and $1.88 per million output tokens, available now via OpenRouter.
Google DeepMind Launches SL2T, a Sign-Language-to-Text Model Trained on 100,000+ Hours Across 50+ Languages
Google DeepMind has released SL2T, a massively multilingual sign-language-to-text translation model trained on over 100,000 hours of data across 50+ sign languages. The model powers new sign-to-text dictation features in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language to English.
Alibaba Releases Qwen3.8 Max, a Multimodal Reasoning Model with 1M Token Context
Alibaba has moved Qwen3.8 Max out of preview into general availability, positioning it as the flagship of the Qwen3.8 series with a 1 million token context window and multimodal input support. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens via OpenRouter.
MiniMax H3 Becomes First Open Video Model to Top an AI Video Ranking
MiniMax has released open weights for H3, a 33-billion-parameter video model that ranks first in Video Editing and second in Text-to-Video on Artificial Analysis — the first time an open model has topped a video generation category. The model accepts text, images, video, and audio in a single prompt, though its highest-resolution module remains closed.
MiniMax Releases H3, a 33B-Parameter Omni-Modal Model That Generates 2K Video With Native Stereo Audio
MiniMax has published MiniMax-H3, a 33-billion-parameter omni-modal generative model capable of producing up to 15 seconds of 2K video with native stereo audio. The model accepts text, image, video, and audio inputs, though its full 2K pipeline depends on a hosted preprocessing component not included in the open-source release.
ByteDance's Seedance 2.5 Generates 30-Second AI Video Clips With Synced Audio
ByteDance released Seedance 2.5, an AI video model that generates synchronized video and audio in a single pass, producing clips up to 30 seconds long that can be extended further. That's roughly triple the length of Google's Gemini Omni Flash.
Moonshot AI Releases Kimi K3: 2.8T-Parameter Open-Weight Model with 1M-Token Context, Now Available via Unsloth Quantiza
Moonshot AI has released Kimi K3, a 2.8-trillion-parameter open-weight mixture-of-experts model with a 1-million-token context window and native multimodal support. Unsloth has published Dynamic 2.0 quantized versions on Hugging Face, claiming improved accuracy over other quantization methods.
Alibaba Launches Qwen3.7 Flash: 1M-Context Vision-Language Model at $0.03/$0.13 per 1M Tokens
Alibaba has released Qwen3.7 Flash, a vision-language reasoning model with a 1 million token context window aimed at multimodal agents, visual coding, and computer-use tasks. The model is priced at $0.03 per 1M input tokens and $0.13 per 1M output tokens and is available through OpenRouter.
Guardoc Health Cuts Documentation Errors 46% Using Amazon Nova Models on Bedrock
Guardoc Health built a multi-stage document processing pipeline on Amazon Nova Pro, Nova Lite, and Titan Text Embeddings to extract and classify medical conditions from clinical PDFs at scale. The company claims a 46 percent reduction in documentation errors, 70 percent fewer audit fines, and over $400K in annual ROI for a single facility.
Moonshot AI Releases Kimi K3: Open-Weight 2.8T-Parameter Model With 1M-Token Context and Native Multimodality
Moonshot AI has released Kimi K3, an open-weight 2.8-trillion-parameter mixture-of-experts model with 104B activated parameters, a 1,048,576-token context window, and native multimodal support. The company describes it as the world's first open 3T-class model, built on a new Kimi Delta Attention architecture.
Black Forest Labs Releases Flux 3, Its First Model to Generate Video With Native Audio Up to 20 Seconds
Black Forest Labs has released Flux 3, a multimodal foundation model trained jointly on images, video, and audio that generates videos up to 20 seconds long with synchronized native audio. The company also introduced Flux-mimic, a robotics action model already being tested at Audi.
Alibaba Releases Qwen-Image-3.0, an Image Generator That Renders 10-Pixel Text and 3x3 Infographic Grids in One Pass
Alibaba's Qwen team has released Qwen-Image-3.0, an image generator that accepts prompts up to 4,500 tokens and can render legible text as small as ten pixels, complex LaTeX formulas, and twelve languages in a single pass. The model is currently invite-only via API, and unlike its predecessor, it likely won't ship with open weights.
Meta launches Muse Image, a free AI image generator integrated across Instagram, WhatsApp, and Facebook Marketplace
Meta has launched Muse Image, a new AI image generator from its Meta Superintelligence Labs division. The model is available free for Instagram Stories, WhatsApp, and the Meta AI app, with integration into Facebook Marketplace for visualizing used furniture in home settings.
Stability AI Releases Stable Audio 3 Medium: 2B-Parameter Audio Generation Model with 180-Second Output in Under 2 Secon
Stability AI has released Stable Audio 3 Medium, a 2 billion parameter latent diffusion model capable of generating variable-length audio up to 380 seconds. The model generates music and sound effects in less than 2 seconds on an H200 GPU, trained on 1.28 million licensed and Creative Commons audio recordings.
Google releases Gemini 3.5 Flash and autonomous agent Gemini Spark at I/O 2026
Google announced Gemini 3.5 Flash and Gemini Spark at I/O 2026. Gemini 3.5 Flash now powers Google's AI Mode search, while Spark is a cloud-based autonomous agent that can monitor credit card statements, track emails, and interact with third-party services like OpenTable and Instacart.
Meta AI app adds voice conversations and live camera features powered by Muse Spark
Meta has added voice conversation capabilities and live camera analysis to its Meta AI app, both powered by the company's Muse Spark model. The features previously available only on Meta's AI glasses are now rolling out to the standalone app, which ranks fourth on the U.S. iPhone App Store.
Amazon Nova 2 Sonic Unifies Speech Recognition, Reasoning, and TTS in Single Streaming Model
Amazon Web Services released technical guidance for migrating text agents to voice assistants using Amazon Nova 2 Sonic, a native speech-to-speech model that combines automatic speech recognition, reasoning, tool calling, and text-to-speech in a single bidirectional streaming interface. The model supports asynchronous tool calling and built-in voice activity detection for handling interruptions.
OpenAI releases ChatGPT Images 2.0 with accurate text rendering and brand-style matching
OpenAI launched ChatGPT Images 2.0, upgrading from decorative images to full-page graphics with detailed text rendering. The update is available to all ChatGPT tiers, with advanced features requiring paid subscriptions that access the Thinking model. Hands-on testing shows significant improvements in text accuracy and brand-style replication, though factual errors still occur.
OpenAI releases ChatGPT Images 2.0 with integrated reasoning and text-image composition
OpenAI has released ChatGPT Images 2.0, which integrates reasoning capabilities to generate complex visual compositions combining text and images. The model supports aspect ratios from 3:1 to 1:3 and outputs up to 2K resolution, with advanced features available to Plus, Pro, Business, and Enterprise users.
Meta launches proprietary Muse Spark model, abandoning open-source commitment
Meta has released Muse Spark, a proprietary AI model with restricted access via API and portal invite only—a striking reversal from CEO Mark Zuckerberg's 2024 manifesto championing open-source AI. The model claims performance matching top competitors from OpenAI, Anthropic, and Google, trained with an order of magnitude less compute than Llama 4.