生成式影片
Wan3.0 Officially Launches: Generate 30-Second Videos with Audio in a Single Run, Using Documents and Web Pages as Direct Inputs
Alibaba has officially launched Wan3.0 following its invite-only testing in early August, unifying support for text, images, video, audio, documents, and web-based references. The API is now public, but the model remains labeled invite-only, with no weights, architecture details, or independent evaluations released.

Alibaba officially launched its Wan3.0 video generation model on August 24, following the public testing that began on August 6. It consolidates text-to-video, first-frame or first-and-last-frame generation, and reference-material-based generation into a single `wan3.0-video` API. It supports output at up to 1080P, generates clips as long as 30 seconds in a single run, and can produce synchronized audio. Unlike previous workflows that primarily accepted prompts and media assets, the new version can directly read DOC, XLS, PPT, PDF, Markdown, Apple Pages, Numbers, and other document formats. It can also parse publicly accessible web pages that do not require authentication. Each request is limited to one document or URL, files must not exceed 100MB, and some formats are capped at 50 pages.
The engineering interface uses an asynchronous task model: users first submit a task to either the Beijing or Singapore endpoint, then poll for the result using a `task_id` that remains valid for 24 hours. The model, API key, and endpoint must all belong to the same region. According to Alibaba, generation typically takes about one to five minutes. Pricing in the Beijing region is calculated per second of output: RMB 0.3 for 480P, RMB 0.6 for 720P, and RMB 1.2 for 1080P. At list price, a full 30-second 1080P video therefore costs RMB 36. For content pipelines, the real innovation is not simply the longer clip duration, but the model’s ability to turn presentation structures, spreadsheet data, or web content into storyboards, animated charts, and voice-overs, reducing the need for preprocessing and manual rewriting.
However, an “official launch” does not mean the model is open. The official catalog still labels the service as invite-only, while Wan’s official GitHub repository and Hugging Face organization provide no Wan3.0 weights, inference code, or license. Details of the architecture, training data, and reproducible benchmarks are also unavailable. Claims of improved character consistency, micro-expressions, and object stability currently rely mainly on vendor demonstrations, and Alibaba acknowledges that audio quality and the accuracy of text rendered in video still need improvement. Developers should next watch for general availability, real-world queue latency, document-parsing error rates, and whether identity and scene drift are genuinely controlled across clips lasting up to 30 seconds.