Alibaba launches Wan 3.0 video model with 30-second single-pass generation and native audio
Alibaba (Tongyi Lab)
Alibaba Cloud released Wan 3.0 on 2026-08-24 as a callable API on Model Studio / Qwen Cloud and the create.wan.video playground. The model generates up to 30 seconds in a single pass (double the 15-second ceiling of Wan 2.7), outputs native 1080p video with audio rendered in the same pass, and accepts text, image, video, audio and document (PDF / slide / web page) inputs. A 'thinking' mode is required for document inputs and lets the model reason about composition before rendering. Three endpoints ship at launch — T2V, I2V (with optional end-frame control), and reference-to-video — accepting up to 10 reference images, 5 reference video clips (15 s total at 16 fps+) and 5 reference audio tracks (15 s total). Pricing is $0.05/$0.10/$0.20 per second at 480p/720p/1080p, with a 'Prime' tier at $0.068/$0.14/$0.28. A public beta had been running since 2026-08-06; the Monday announcement coincided with the close of Alibaba's HK$80 billion follow-on share sale — the largest in Hong Kong since 2021.
Why it matters
Wan 3.0 doubles the single-pass video length ceiling set by the prior Wan generation and adds same-pass native audio, bringing Alibaba's flagship video model into the same feature bracket as ByteDance Seedance 2.5 and Hailuo H3. The reference composition (images / video clips / audio clips / documents as first-class inputs) positions it as a unified multimodal model rather than a pure text-to-video system. Wan remains API-only — Wan 2.2 is still the last open-weights flagship, so for open-source consumers this is not yet a drop-in replacement.
Importance: 4/5
Frontier-tier video model release (Alibaba Wan 3.0)