model release

Alibaba Launches Qwen-Audio-3.1, Cuts AI Audio API Prices by Up to 95 Percent

TL;DR

Alibaba's Qwen team has released Qwen-Audio-3.1, a five-model lineup covering speech recognition, text-to-speech, and real-time voice interaction. Alongside the release, Alibaba cut API pricing by up to 95 percent for ASR, 85 percent for real-time models, and 70 percent for TTS.

2 min read
0

Alibaba's Qwen team has released Qwen-Audio-3.1, a five-model family spanning automatic speech recognition (ASR), text-to-speech (TTS), and real-time voice interaction. The release comes with steep API price cuts: up to 95 percent for ASR, roughly 85 percent for the real-time model, and about 70 percent for TTS, according to Alibaba.

What's in the lineup

The base ASR model improves multilingual and dialect recognition and automatically strips filler words and repeated phrases from transcripts. ASR-Next extends this with multi-speaker identification and timestamps, plus detection of emotional tone, ambient sounds, and machine noise — useful for transcribing meetings, calls, or noisy field recordings.

On the generation side, TTS supports multilingual synthesis with cross-language voice transfer, meaning a cloned voice can speak in a different language than its source recording. Users steer emotion, pacing, and style through natural-language prompts — Alibaba's example is instructing the model to "read this with a sharp, commanding tone, demanding respect." TTS-Next combines a language model with a diffusion-based audio generator to produce voice, sound effects, and background audio in a single generation pass, rather than layering them separately.

The fifth model, Realtime, handles simultaneous speaking and listening with instant interruption support — a requirement for natural voice-agent conversations where users cut in mid-response. According to Qwen, the model also detects low mood in a speaker's voice and adjusts its own delivery to speak more slowly and empathetically.

Pricing and availability

Alibaba has not published exact per-token or per-minute rates for each model in the materials reviewed. The company states discounts of up to 95 percent for ASR, approximately 85 percent for Realtime, and about 70 percent for TTS relative to prior Qwen audio pricing. Full pricing tables and API documentation are available through Alibaba's Qwen Cloud console and blog; specific dollar figures were not disclosed in the announcement.

No formal benchmark scores accompanied the release. Claims about improved dialect recognition, emotion detection accuracy, and voice-transfer quality come from Alibaba and have not been independently verified.

What this means

This release targets the voice-agent and transcription infrastructure layer rather than chasing headline benchmark wins. Multi-speaker diarization with emotion and noise detection (ASR-Next) and single-pass generation of voice plus sound effects (TTS-Next) are features aimed squarely at companies building call-center automation, dubbing pipelines, and conversational agents — markets where OpenAI's Realtime API, ElevenLabs, and Google's audio models already compete.

The price cuts matter more than the feature list in the near term. A 95 percent reduction in ASR costs, if it holds up against usage-based billing at scale, meaningfully changes the economics of high-volume transcription products and could pressure competitors to follow with their own cuts. Whether Qwen-Audio-3.1's quality holds up against incumbents in independent testing — rather than Alibaba's own descriptions — remains to be seen, since no third-party benchmark data was published alongside the launch.

Related Articles

model release

NVIDIA Releases Nemotron 3 Diarization, a 100M-Parameter Open-Weight Model Ranked #1 on Voice Arena's Diarization-Bench

NVIDIA released Nemotron 3 Diarization, a 100-million-parameter open-weight model that identifies who is speaking and when in audio conversations. It ranked #1 among 17 system configurations on Voice Arena's Diarization-Bench with a 14.72% diarization error rate, supporting up to eight speakers in both live and recorded audio.

model release

Qwen3.8 Omni Flash: Alibaba's First Agentic Omni-Modal Model Adds Native Audio-Video Understanding, 1M Context

Alibaba's Qwen team has released Qwen3.8 Omni Flash, described as the first Qwen model built around agentic capabilities with native audio-video understanding. It ships with a 1M-token context window and support for two- and four-channel spatial audio.

model release

OpenAI Launches GPT-6 Sol Pro, a High-Reasoning Mode for Its Mid-Tier Model at $2/$10 per 1M Tokens

OpenAI has released GPT-6 Sol Pro, which runs the existing GPT-6 Sol model with its reasoning mode set to 'pro' for higher-quality responses on complex tasks. It carries a 1.1 million token context window and is priced at $2 per 1M input tokens and $10 per 1M output tokens.

model release

OpenAI Launches GPT-6 Sol and GPT-6 Luna, Cutting API Prices 50% Versus GPT-5.6

OpenAI has released GPT-6 Sol and GPT-6 Luna, two new models that cost 50% less than their GPT-5.6 equivalents while claiming improved coding and computer-use performance. The models roll out today to ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users.

Comments

Loading...