Alibaba's Wan3.0 Breaks the 30-Second Video Barrier With Document-to-Video

Alibaba's Wan3.0 enters public beta with 30-second native video generation, document-to-video inputs, and API pricing starting at $0.05/sec

ByWanWan
·
·
AuthorWan
Read1 min
TopicVideo · Api
  • Alibaba's Wan3.0 enters public beta on Qwen Cloud and Alibaba Cloud Model Studio, with full API rollout coming soon.
  • Native 30-second single-pass video generation, up from 2-15 seconds in Wan 2.7, eliminating the need to stitch clips together.
  • Omni-Reference accepts documents, spreadsheets, slides, PDFs, and live URLs as creative inputs alongside text, image, audio, and video.
  • API pricing: $0.05/sec at 480p, $0.10/sec at 720p, $0.20/sec at 1080p -- a full 30-second 1080p clip costs $6.
  • Built on a Diffusion Transformer (DiT) architecture with production-grade character consistency and in-pass audio generation.
  • Community is cautiously excited but watching whether Alibaba will release open weights under Apache 2.0, given a history of broken open-source promises.

Alibaba just opened public beta access to Wan3.0, its most capable video generation model yet. The headline features are hard to ignore: native 30-second video clips in a single generation pass, what Alibaba calls "reality-grade rendering," and a genuinely novel input system that goes well beyond the text-and-image prompts every other video model accepts.

The 30-second barrier, broken

Until now, most AI video models topped out at short clips. Early testers described roughly 30 seconds of native single-shot video from Wan3.0, compared to Wan 2.7 which was documented at 2 to 15 seconds. That jump matters more than it sounds. A 30-second clip covers a full broadcast spot in one generation, with no cuts to stitch. The previous workflow of generating multiple short clips and editing them together is now optional, not mandatory.

Omni-Reference: the real differentiator

The feature that sets Wan3.0 furthest apart from competitors is what Alibaba calls Omni-Reference. Every other video model takes text, maybe an image, and occasionally audio. Wan3.0 takes all of that, plus structured documents.

Supported input formats include:

  • Documents: .doc, .xls, .ppt, .pdf, .txt, .key, .pages, .numbers

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves