model release

Alibaba Launches Wan3.0, Generating AI Videos Up to 30 Seconds From Text, Images, and Documents

TL;DR

Alibaba's Wan3.0 video generation model is now in beta, producing clips up to 30 seconds long and accepting text, images, video, audio, and documents like PDFs and PowerPoint files as input. Pricing runs per second of output across two tiers and three resolutions.

2 min read
1

Alibaba has launched Wan3.0, a new video generation model now available in beta, capable of producing clips up to 30 seconds long — double the maximum length of its predecessor, Wan2.5. The model accepts an unusually broad range of inputs: text, images, video, audio, and documents including PDFs, PowerPoint files, and web pages.

A single prompt can combine up to ten images, five videos, and five audio clips simultaneously. According to Alibaba, this lets users turn static source material — a slide deck, a webpage, a set of product photos — directly into video content. Wan3.0 also recommends an optimal clip length based on the prompt and includes a tool to extend existing videos beyond their original runtime.

Alibaba claims the model addresses a common failure mode in AI video generation: visual drift, where faces, interfaces, and objects distort or lose consistency over the course of a clip. The company says Wan3.0 preserves details from reference material — characters, props, spatial layouts — more reliably than prior versions, though this claim has not been independently verified.

Pricing and availability

Wan3.0 is accessible through the wan.video website, Alibaba Cloud Model Studio, and via API on Qwen Cloud. It ships in two tiers: a Standard version (currently discounted 30 percent) and a faster Prime version. Pricing is billed per second of generated video and scales with resolution:

Resolution Standard (per sec.) Prime (per sec.) 30-sec clip (Standard) 30-sec clip (Prime)
480p $0.05 $0.068 $1.50 $2.04
720p $0.10 $0.14 $3.00 $4.20
1080p $0.20 $0.28 $6.00 $8.40

A full 30-second 1080p clip on the Prime tier costs $8.40 before any promotional discount.

Target use cases

Alibaba is positioning Wan3.0 across a wide set of markets: speeding up film production, generating short-form drama and social media content, producing marketing and training videos from existing business documents, and creating simulation footage for training autonomous vehicles and robotics systems.

The launch lands as Alibaba sharply increases AI capital expenditure. The company recently completed the largest share sale by a Hong Kong-listed firm specifically to fund its AI buildout, and last week reported a 75 percent year-over-year drop in quarterly profit, which it attributed to higher AI investment spending.

What this means

Wan3.0's document-to-video capability — turning PDFs and slide decks directly into motion content — is a distinct product angle from competitors like Runway, Pika, and OpenAI's Sora, which lean primarily on text and image prompts. Combined with per-second pricing that undercuts many rivals' per-clip rates at lower resolutions, Alibaba is clearly chasing enterprise and content-production workflows rather than just prosumer creative tools. Whether the claimed consistency improvements hold up against independent benchmarks — and whether 30-second generation quality degrades toward the end of clips, a common issue in longer AI video — remains to be tested outside Alibaba's own demos.

Related Articles

model release

Alibaba's Qwen Releases Qwen-Drive-1.0-4B, a Unified VLM for Autonomous Driving Perception and Planning

Alibaba's Qwen team has released Qwen-Drive-1.0-4B, a 4B-parameter vision-language model built on Qwen3.5 that unifies 3D perception, driving question answering, and motion planning in one framework. The model reports strong open-loop, pseudo-closed-loop, and closed-loop driving benchmark results while claiming minimal loss of general vision-language ability.

model release

Alibaba Open-Sources Qwen3.8-2.4T-A95B, Its First Qwen-Max-Class Model With Public Weights

Alibaba's Qwen team released Qwen3.8-2.4T-A95B on August 12, 2026, the open-weight version of Qwen3.8-Max and the first Qwen-Max-class model made publicly available. The 2.4 trillion-parameter mixture-of-experts model activates only 95 billion parameters per token and supports context windows up to 1 million tokens.

model release

AllSpark's Iris-mini and Iris-pro Top Open-Weight Search Agent Benchmarks

Chinese lab AllSpark has released Iris-mini and Iris-pro, two open-weight search agents built on Qwen3 models that claim the top spot among open-weight systems in their size classes on four research benchmarks. The release includes model weights, an agent harness, and evaluation code, with training pipelines to follow.

model release

Tencent Open-Sources AuK, a 1.5B-Parameter Speech Generation and Editing Model

Tencent has open-sourced AuK, a 1.5B-parameter foundation model for speech generation and editing that handles TTS, content editing, and audio enhancement through natural-language instructions. The release includes a distilled AuK-Flash variant for 4-step fast inference, both under MIT license.

Comments

Loading...

Alibaba Wan3.0: 30-Second AI Video Generator Launches in Beta | TPS