Alibaba Launches Wan3.0, Generating AI Videos Up to 30 Seconds From Text, Images, and Documents
Alibaba's Wan3.0 video generation model is now in beta, producing clips up to 30 seconds long and accepting text, images, video, audio, and documents like PDFs and PowerPoint files as input. Pricing runs per second of output across two tiers and three resolutions.
Alibaba has launched Wan3.0, a new video generation model now available in beta, capable of producing clips up to 30 seconds long — double the maximum length of its predecessor, Wan2.5. The model accepts an unusually broad range of inputs: text, images, video, audio, and documents including PDFs, PowerPoint files, and web pages.
A single prompt can combine up to ten images, five videos, and five audio clips simultaneously. According to Alibaba, this lets users turn static source material — a slide deck, a webpage, a set of product photos — directly into video content. Wan3.0 also recommends an optimal clip length based on the prompt and includes a tool to extend existing videos beyond their original runtime.
Alibaba claims the model addresses a common failure mode in AI video generation: visual drift, where faces, interfaces, and objects distort or lose consistency over the course of a clip. The company says Wan3.0 preserves details from reference material — characters, props, spatial layouts — more reliably than prior versions, though this claim has not been independently verified.
Pricing and availability
Wan3.0 is accessible through the wan.video website, Alibaba Cloud Model Studio, and via API on Qwen Cloud. It ships in two tiers: a Standard version (currently discounted 30 percent) and a faster Prime version. Pricing is billed per second of generated video and scales with resolution:
| Resolution | Standard (per sec.) | Prime (per sec.) | 30-sec clip (Standard) | 30-sec clip (Prime) |
|---|---|---|---|---|
| 480p | $0.05 | $0.068 | $1.50 | $2.04 |
| 720p | $0.10 | $0.14 | $3.00 | $4.20 |
| 1080p | $0.20 | $0.28 | $6.00 | $8.40 |
A full 30-second 1080p clip on the Prime tier costs $8.40 before any promotional discount.
Target use cases
Alibaba is positioning Wan3.0 across a wide set of markets: speeding up film production, generating short-form drama and social media content, producing marketing and training videos from existing business documents, and creating simulation footage for training autonomous vehicles and robotics systems.
The launch lands as Alibaba sharply increases AI capital expenditure. The company recently completed the largest share sale by a Hong Kong-listed firm specifically to fund its AI buildout, and last week reported a 75 percent year-over-year drop in quarterly profit, which it attributed to higher AI investment spending.
What this means
Wan3.0's document-to-video capability — turning PDFs and slide decks directly into motion content — is a distinct product angle from competitors like Runway, Pika, and OpenAI's Sora, which lean primarily on text and image prompts. Combined with per-second pricing that undercuts many rivals' per-clip rates at lower resolutions, Alibaba is clearly chasing enterprise and content-production workflows rather than just prosumer creative tools. Whether the claimed consistency improvements hold up against independent benchmarks — and whether 30-second generation quality degrades toward the end of clips, a common issue in longer AI video — remains to be tested outside Alibaba's own demos.
Related Articles
Qwen 3.8 27B Launches with Vision Support and a 262K Context Window—But Its Default Settings Cause Massive Overthinking
Alibaba's Qwen research lab has released Qwen 3.8 27B, an Apache 2.0 licensed, vision-capable model with a 262,144-token context window. Independent testing found the model's default 'xhigh' reasoning setting causes it to massively overthink simple prompts, turning quick tasks into 20-minute ordeals.
Alibaba Releases Qwen3.8 Open-Weight Models Under Apache 2.0, Including 27B Multimodal Model with 262K Native Context
Alibaba's Qwen team has released open weights for Qwen3.8, including a 27-billion-parameter multimodal dense model with 262,000 tokens of native context. The models ship under the Apache 2.0 license and are available on Hugging Face and ModelScope.
Alibaba Releases Qwen3.8-27B-FP8, a 27B Dense Vision-Language Model with 1M-Token Context
Alibaba's Qwen team has released FP8-quantized weights for Qwen3.8-27B, a 27-billion-parameter dense vision-language model with native 262,144-token context extensible to 1 million tokens. The model claims gains over its Qwen3.6 and Qwen3.7 predecessors on coding, agentic, and multimodal benchmarks.
Alibaba Releases Qwen3.8-27B, a Dense Vision-Language Model with 1M-Token Context
Alibaba's Qwen team has released Qwen3.8-27B, a 27-billion-parameter dense vision-language model with 262,144-token native context extensible to 1 million tokens. The model shows gains over Qwen3.6-27B and Qwen3.7-Plus across coding, agentic, and multimodal benchmarks, according to Alibaba.
Comments
Loading...