
Alibaba flipped Wan 3.0 from public beta to GA on Monday, and the thing that should stop you cold isn't the 30-second one-take generation — it's the part where you hand the model a PDF and it spits out a thirty-second video without a script, an editor, or a human in the loop. That's the line that crossed.
Wan 3.0 takes text, images, video, audio, and now PDFs, DOC, XLS, PPT, and MD files (up to 100 MB or 50 pages) plus web pages by URL, and produces a finished 1080p video in a single inference pass. No competitor — not Veo, not Kling, not Seedance, not Gemini Omni Flash — accepts documents as a first-class input. Alibaba shipped something nobody else has, and most of the coverage missed it because they got stuck on the duration number. (Reuters, Alibaba Cloud Model Studio)
Previous Wan versions topped out at 15 seconds per generation. Wan 3.0 doubles that to 30 seconds in a single pass — no stitching of short fragments, no temporal artifacts at the seams. Audio is rendered in the same forward pass as the video, locked to frame one, so dialogue and ambient sound stay in sync across the whole take.
That changes the production math in a specific way: a continuous one-take language — the dolly shot, the walk-and-talk, the unbroken Steadicam move — only becomes possible when the model can hold coherence across the whole duration. Before Wan 3.0, you either accepted a montage of 6-second cuts or you paid for an editor to glue shots together. After it, the editor's job shifts from assembly to direction. That's a real workflow change, not a benchmark bragging point.
The model also picks its own duration now. You describe a scene and it proposes how long the clip should be rather than you guessing. Magnific's team used this to ship an entire short film on Wan 3.0 with thirteen scenes and locked character plates across three ages per character. That's not a tech demo — that's production work. (Magnific)
Drop a quarterly report into the model and you get a thirty-second video summary. Feed it a slide deck and you get a narrated walkthrough. Point it at a PDF manual and you get training footage. That's the pitch Alibaba is making and, watching the demos, it's not vaporware — the model is clearly grounding the generated scenes in the actual document structure, not hallucinating content unrelated to the input.
For enterprise teams who have been quietly scripting training videos, product demos, and explainer content by hand for years, Wan 3.0 removes the writer step entirely. You feed it the source material; it writes, directs, animates, and scores the video in one call. Audio quality and on-screen text rendering still have rough edges — Alibaba is honest about that in their own materials — but the rough edges are improving fast. The fact that it exists, from a team with Alibaba's distribution, changes the conversation permanently.
Here's the part Alibaba doesn't want you to notice: Wan 3.0 is closed. The last open flagship of the Wan line was Wan 2.2, shipped under Apache 2.0 in summer 2025. The community has been reminding Alibaba of its broken promise to open-source 2.5 ever since, and 3.0 doubles down on that retreat. Search results are already full of fake "Wan 3.0 free download" and "Apache 2.0" affiliate sites. None of those are real. The weights are closed, the API is paid, and the international rollout is gradual.
I get why Alibaba did it — the model is good enough to be a product, and a paid product funds the next training run. But the strategic cost is real. The open-weights Qwen 3.8 Max and 27B Apache release earlier this month gave Alibaba genuine developer mindshare in the LLM world. Closing Wan 3.0 surrenders that exact advantage in the video category to ByteDance, which is still publishing Seedance weights more aggressively, and to the open community that's been fine-tuning Wan 2.2 for two years.
The pricing is straightforward: $0.05, $0.10, and $0.20 per second of output at 480p, 720p, and 1080p respectively. A thirty-second 1080p clip runs about $6. That's roughly a third more than Wan 2.7, which is reasonable given the doubled duration and the unified model surface. It's also the first video model pricing that starts to feel like infrastructure pricing rather than demo pricing — you're not paying per experiment, you're paying per finished asset.
The other number is bigger: Alibaba launched a $10 billion share placement the same week to fund AI compute. Wan 3.0 is the first product of that raise hitting the public. Read those two announcements together. This is the moment where the Chinese cloud giants stop running AI as a research line item and start treating it as a capital expenditure category on par with hyperscaler infrastructure spend. (Reuters)
Wan 2.7 holds fourth place on the Video Arena leaderboard (Elo 1161), behind Gemini Omni Flash, Hailuo H3, and Seedance 2.0. Seedance 2.5 hit the 30-second one-take mark a week before Wan 3.0, and ByteDance still wins on raw reference capacity — Seedance accepts up to 50 inputs per call, versus Wan's 15. But Wan 3.0 owns documents and web pages, and the editing feature — regenerate a selected interval, swap dialogue, rewrite plot points without re-rendering the whole clip — is a workflow primitive neither Seedance nor Veo ship today.
Native resolution is 1080p, not 4K. The "Wan 3.0 does 4K" claims in search results are misinformation — use Magnific or ComfyUI upscalers on top if you need 4K deliverables.
The closing take: the video model race just stopped being a duration race. It's now a multimodal-input race, and Alibaba shipped a feature nobody else has. If you're building any content pipeline that touches documents, slides, or spreadsheets, test Wan 3.0 this week. The closed-weights situation means you'll be renting, not owning — but the productivity delta is large enough that renting is the right call until an open-weights competitor shows up with the same document-to-video primitive.