Get the app

Alibaba Ships Wan 3.0: 30-Second Single-Pass Video and Omni-Reference Control

Collapsing fragmented video generation models into a unified omni-modal engine, Wan 3.0 pairs 30-second single takes with native synchronized audio and multi-asset referencing.

Alibaba’s Tongyi Lab has launched Wan 3.0 (wan3.0-video), rolling out a major paradigm shift in generative video across DashScope, Qwen Cloud, and ComfyUI v0.33.4. Rather than forcing creators to juggle fragmented checkpoints for distinct tasks, Wan 3.0 collapses text-to-video, image-to-video, video-to-video editing, and reference steering into a single unified foundation model capable of rendering unbroken 30-second continuous takes with natively co-synthesized audio.

The release addresses the two biggest pain points in AI video production: the temporal decay caused by stitching short 4-to-5-second clips and the brittle conditioning systems that fail to preserve identity and camera kinematics across frames.

Unifying Fragmented Pipelines into Omni-Modal Reference

Previous state-of-the-art architectures, including the Wan 2.7 generation, required developers to route requests across specialized pipelines depending on the task: one model variant for text-to-video (T2V), another for image-to-video (I2V), and separate checkpoints for video-to-video editing or camera trajectory conditioning.

Wan 3.0 eliminates this pipeline fragmentation with an Omni-Reference architecture. The model accepts up to 20 reference assets simultaneously in a single generation request:

  • Up to 10 image references for character sheets, wardrobe specifications, and environmental textures.
  • Up to 5 video references to guide camera trajectories, actor pacing, and spatial motion dynamics.
  • Up to 5 audio references to dictate vocal timbre, background acoustics, and rhythm.
  • Document and URL ingestion, enabling the model to directly parse PDF slide decks, proposal sheets, or live web pages as contextual source material.

Conditioning across these assets is managed through explicit token indexing using an @ syntax (e.g., @Image1, @Video2, @Audio1). A prompt can explicitly instruct the engine: "The character from @Image1 picks up the artifact in @Image2 while the camera tracks the continuous dynamic motion of @Video1, set to the vocal cadence of @Audio1." This fine-grained reference routing ensures that facial geometry, lighting consistency, and camera framing remain locked throughout the scene.

True 30-Second Single-Pass Generation and Native Audio

Most video generators enforce short 5-to-15-second windows because autoregressive extensions and latent chunk stitching suffer from compounding error drift—resulting in face warping, lighting inconsistency, and hallucinatory background artifacts.

Wan 3.0 bypasses clip stitching by rendering up to 30 seconds in a single latent diffusion pass. This unlocks genuine narrative structure: single-take sequence shots, continuous pans, and dynamic action blocks that maintain spatial persistence from frame 1 to frame 720 (at 24 fps).

Key architectural capabilities delivered in this pass include:

  • Joint Audio-Visual Co-Generation: Audio is no longer an afterthought bolted on via downstream text-to-speech or foley post-processing. Wan 3.0 natively synthesizes dialogue, lip synchronization, vocal performances (including singing and rapping), ambient room tone, and action-timed acoustic effects directly inside the primary generation process.
  • Instruction-Guided Latent Video Editing: The model supports granular in-place video transformation. Creators can issue natural language prompts—such as "have the actor put down the mug and step toward the window"—to alter subjects or actions while preserving the rest of the frame's temporal integrity.
  • Sharp Interface & Motion Graphic Rendering: Previous diffusion models routinely garbled fine typographic details and vector-like interface elements. Wan 3.0 introduces high-fidelity digital scene rendering, making UI walkthroughs, screencasts, and typography-heavy marketing assets legible without secondary super-resolution passes.

Pricing Tiers and Ecosystem Availability

Following its initial DashScope integration, Wan 3.0 is live across developer platforms including Qwen Cloud, Fal.ai, Atlas Cloud, and local ComfyUI environments via the newly released ComfyUI v0.33.4 update.

Alibaba has structured the API around resolution-based consumption rather than flat compute tokens:

  • 480P Draft Tier: $0.05 per second of generated video (designed for rapid prototyping, storyboard pre-visualization, and fast testing).
  • 720P Standard Tier: $0.10 per second of generated video.
  • 1080P High-Definition Tier: $0.20 per second of generated video.

The API supports arbitrary aspect ratios—including standard 16:9, 9:16 vertical, 4:3, 3:4, and 1:1 square—alongside an adaptive mode that dynamically sizes the frame based on input reference dimensions. Duration can be set anywhere between 2 and 30 seconds, or assigned to auto to let the model determine optimal shot length based on narrative tempo.

Grounding the Spec: Debunking 4K and Open-Weight Rumors

The arrival of Wan 3.0 triggered widespread speculation across developer channels, with marketing wrappers asserting native 4K output and open weights under Apache 2.0. However, first-party technical documentation and endpoint inspection clarify the current reality:

  • No Native 4K Output: Despite speculative claims, Alibaba’s production endpoints top out at 1080p. There is no 4K billing tier or API flag; 4K outputs currently shared online rely on third-party post-upscaling.
  • API-Only Public Beta: While earlier models like Wan 2.1 and Wan 2.2 were open-weighted under Apache 2.0, Wan 3.0 is currently hosted exclusively as an API service (wan3.0-video). Alibaba has not yet released model weights on Hugging Face or ModelScope.
  • Concurrency and Queueing: Production endpoints operate with a concurrency ceiling of 2 simultaneous requests per developer key, a 50-task asynchronous queue buffer, and a 30 requests-per-minute (RPM) rate limit.

What It Means for AI Media Pipelines

Wan 3.0 represents a clean break from fragmented AI video workflows. By consolidating reference conditioning, native audio synchronization, and 30-second continuous temporal modeling into a single callable endpoint, Alibaba has eliminated several layers of pipeline glue code.

For studio workflows and independent developers, the immediate unlock is practical control: complex character consistency, directed camera moves, and coherent scene-length narratives are now achievable in a single prompt and seed, without the fragile multi-model handoffs that previously defined generative video production.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play