model release

Alibaba Launches Wan3.0, Generating AI Videos Up to 30 Seconds From Text, Images, and Documents

TL;DR

Alibaba's Wan3.0 video generation model is now in beta, producing clips up to 30 seconds long and accepting text, images, video, audio, and documents like PDFs and PowerPoint files as input. Pricing runs per second of output across two tiers and three resolutions.

2 min read
1

Alibaba has launched Wan3.0, a new video generation model now available in beta, capable of producing clips up to 30 seconds long — double the maximum length of its predecessor, Wan2.5. The model accepts an unusually broad range of inputs: text, images, video, audio, and documents including PDFs, PowerPoint files, and web pages.

A single prompt can combine up to ten images, five videos, and five audio clips simultaneously. According to Alibaba, this lets users turn static source material — a slide deck, a webpage, a set of product photos — directly into video content. Wan3.0 also recommends an optimal clip length based on the prompt and includes a tool to extend existing videos beyond their original runtime.

Alibaba claims the model addresses a common failure mode in AI video generation: visual drift, where faces, interfaces, and objects distort or lose consistency over the course of a clip. The company says Wan3.0 preserves details from reference material — characters, props, spatial layouts — more reliably than prior versions, though this claim has not been independently verified.

Pricing and availability

Wan3.0 is accessible through the wan.video website, Alibaba Cloud Model Studio, and via API on Qwen Cloud. It ships in two tiers: a Standard version (currently discounted 30 percent) and a faster Prime version. Pricing is billed per second of generated video and scales with resolution:

Resolution Standard (per sec.) Prime (per sec.) 30-sec clip (Standard) 30-sec clip (Prime)
480p $0.05 $0.068 $1.50 $2.04
720p $0.10 $0.14 $3.00 $4.20
1080p $0.20 $0.28 $6.00 $8.40

A full 30-second 1080p clip on the Prime tier costs $8.40 before any promotional discount.

Target use cases

Alibaba is positioning Wan3.0 across a wide set of markets: speeding up film production, generating short-form drama and social media content, producing marketing and training videos from existing business documents, and creating simulation footage for training autonomous vehicles and robotics systems.

The launch lands as Alibaba sharply increases AI capital expenditure. The company recently completed the largest share sale by a Hong Kong-listed firm specifically to fund its AI buildout, and last week reported a 75 percent year-over-year drop in quarterly profit, which it attributed to higher AI investment spending.

What this means

Wan3.0's document-to-video capability — turning PDFs and slide decks directly into motion content — is a distinct product angle from competitors like Runway, Pika, and OpenAI's Sora, which lean primarily on text and image prompts. Combined with per-second pricing that undercuts many rivals' per-clip rates at lower resolutions, Alibaba is clearly chasing enterprise and content-production workflows rather than just prosumer creative tools. Whether the claimed consistency improvements hold up against independent benchmarks — and whether 30-second generation quality degrades toward the end of clips, a common issue in longer AI video — remains to be tested outside Alibaba's own demos.

Related Articles

model release

Reka AI releases Rho-1, a 19B-parameter omni-model for text, image, video and robot control

Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single network. Reka says it uses no tool calls or external models. Context window, pricing, and benchmark scores have not been disclosed.

model release

StepFun releases Step 5 Preview: 600B MoE with 1M context at $1/$2.70 per 1M tokens

StepFun has listed Step 5 Preview, a sparse Mixture-of-Experts model with 600B total and 27B active parameters and a 1.0M-token context window. It is priced at $1 input and $2.70 output per 1M tokens on OpenRouter. StepFun positions it as its flagship model for agentic work.

model release

OpenAI launches GPT-6 in ChatGPT with 'Intelligent UI' and interactive answers; Sol for paid users, Luna for free

OpenAI is rolling out GPT-6 to all ChatGPT tiers, with paying users on GPT-6 Sol and free users on GPT-6 Luna. The release adds 'Intelligent UI,' which renders answers as interactive charts, buttons, forms and mini apps, and lets the model respond while still thinking. OpenAI claims this cuts wait times by 44 percent.

model release

Claude Haiku 5.5 arrives on Amazon Bedrock; Anthropic claims ~75% lower cost than Haiku 4.5

Claude Haiku 5.5 is available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, it is the fastest and most efficient model in the Claude 5.5 family and costs around 75% less than Claude Haiku 4.5 for most tasks. It is the first Haiku model with effort controls.

Comments

Loading...