Alibaba has launched Wan3.0, the latest version of its video generation model, days after raising roughly $10.2bn in a Hong Kong share placement whose proceeds it has earmarked entirely for AI. The sequencing is not subtle, and it is not meant to be, since the company has spent the past fortnight arguing to investors that its capital expenditure is buying something they can see.
The model had already been running in public beta since early August through Alibaba Cloud’s Model Studio and its Qwen Cloud platform, so Monday’s launch is a widening of access rather than a first appearance. It also lands in a category where Chinese labs have quietly built a commanding position while Western attention has stayed on chatbots.
Wan3.0 generates clips of up to 30 seconds in a single pass, double the 15-second ceiling of its predecessor Wan2.7, at resolutions up to 1080p. Alibaba says the model holds character detail, props, spatial layout, and motion graphics steady across the full length and renders faces with synchronised micro-expressions and multilingual voice output.
The more distinctive feature is what it will accept as input. Alongside text, images, audio, and video, Wan3.0 takes web pages and documents, including PDFs and PowerPoint files, which lets a user hand it a slide deck or a spreadsheet and get back a video that reflects the structure of the source material.










