model release

Qwen3.8-Omni-Flash Prices Multimodal AI at $0.15/$0.47 per Million Tokens, Undercutting Gemini Flash by 5x

TL;DR

Alibaba's Qwen team released Qwen3.8-Omni-Flash, a multimodal model for AI agents that processes audio and video with a 1 million token context window. Pricing undercuts Google's Gemini 3.8 Flash by roughly 5x on input and 8x on output, according to Qwen.

2 min read
0

Qwen3.8-Omni-Flash — Quick Specs

Context window1000K tokens
Input$0.15/1M tokens
Output$0.47/1M tokens

What happened

Alibaba's Qwen team has released Qwen3.8-Omni-Flash, described as the company's first multimodal model built specifically for AI agents. The model processes audio and video simultaneously, draws conclusions from combined inputs, and can autonomously invoke tools to edit vlogs, translate short videos, or summarize movies.

The context window spans 1 million tokens. According to Qwen, the model performs close to Google's Gemini 3.8 Flash on audio-video benchmarks — though the company has not published specific comparative benchmark scores alongside the claim.

Pricing

Qwen3.8-Omni-Flash is priced at $0.15 per million input tokens and $0.47 per million output tokens through the Qwen API. Qwen estimates audio input costs under $0.01 per hour of content, and processing 720p video with audio at one frame per second runs about $0.20, excluding response generation costs.

By comparison, Gemini 3.8 Flash charges $0.75 per million input tokens and $3.75 per million output tokens at its introductory rate — roughly 5x and 8x more than Qwen's rates, respectively. Google's pricing is set to double on January 1, 2027.

Availability and tooling

The model is accessible through Qwen Studio, Qwen Cloud, and the standard API. Alongside the model release, Qwen introduced open-source Qwen-MM-Plugins, which add video editing, speaker recognition, PDF-to-video note generation, and reusable agent workflows. These plugins integrate with existing agent frameworks including Claude Code, Gemini CLI, and Qwen Code.

A separate tool, Qwen-Live Harness, enables real-time interaction using a device's camera and microphone, positioning the release toward live agentic use cases rather than purely batch processing.

What this means

Qwen3.8-Omni-Flash is a direct pricing challenge in the multimodal agent space, where Gemini Flash has held a strong position for cost-sensitive, high-volume audio and video workloads. If Qwen's benchmark parity claims hold up under independent testing, the price gap — roughly 5x cheaper on input, 8x cheaper on output — could pull developers building video-processing agents, vlog editors, and real-time assistants toward Alibaba's stack, particularly in cost-per-hour-of-video use cases like the ~$0.20 video-processing estimate.

The open-sourcing of Qwen-MM-Plugins for third-party agent frameworks (Claude Code, Gemini CLI) is notable: it signals Qwen is optimizing for adoption within existing developer tooling rather than trying to build a closed ecosystem. That said, Qwen's benchmark comparisons are self-reported, and Gemini 3.8 Flash's actual real-world audio-video accuracy under production loads hasn't been independently verified against these figures. Buyers evaluating a switch should run their own comparative tests before committing to either platform's pricing model, especially given Google's scheduled price increase in 2027.

Related Articles

model release

PrismML's Bonsai 2 Compresses 27B-Parameter Model to 5.9GB, Retains 98% of Benchmark Performance

PrismML released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B model that shrinks memory footprint by 9x to 10x down to 5.9GB. The startup claims 98% aggregate benchmark parity with the original, up from 95% in its first release, using a ternary weight compression technique.

model release

OpenAI's GPT-6 Astra Beats Pokémon in 18 Hours, Scores 62.7% on ARC-AGI-3

GPT-6 Astra completed Pokémon FireRed in 18 hours 12 minutes, five times faster than its predecessor, and scored 62.7% on ARC-AGI-3 versus 7.78% for GPT-5.6 Sol. The model also ran a 141-hour Minecraft session and finished Fallout 3 in roughly 59 hours, according to independent testers.

model release

Anonymous Provider Launches Union Alpha, a Free 262K-Context Multimodal Model on OpenRouter

A third-party provider using the alias 'Stealth' has released Union Alpha on OpenRouter, a multimodal model with a 262K context window, currently free to use during its preview period. The model's developer remains anonymous, and OpenRouter states it is not the model's owner or operator.

model release

PrismML Releases Ternary Bonsai 2 27B, a Compressed Reasoning Model with 262K Context

PrismML has released Ternary Bonsai 2 27B, a 27B-parameter reasoning model derived from Qwen3.8-27B that uses ternary weight compression to shrink to roughly 8.5 GB. The model supports a 262K-token context window, image understanding, tool calling, and thinks by default at 'xhigh' reasoning effort.

Comments

Loading...