Wan

High-quality AI video and image generation model family built for practical, reliable, and creative workflows.

High-quality AI video generation

Wan is Alibaba’s next generation family of visual generation models, designed to turn text, images, audio, and reference clips into high quality, cinematic outputs.

Wan video generation

Performance & practical strengths

Cinematic-quality output

Wan generates native 1080p videos with improved lighting, texture detail, and shot-to-shot stability. This architecture produces smoother motion, more detailed environments, and a consistent visual style across longer sequences.

Audio–video synchronisation

The models include built-in audio-video synchronisation, ensuring that speech, sound effects, and background audio closely match on-screen movement. This supports realistic lip-sync, expressive character performances, and polished dialogue-driven scenes.

Stable storytelling tools

Shot scheduling and multi-lens narrative control allow users to create multi-scene sequences with smooth transitions and consistent framing. These tools are particularly suited for advertisements, drama clips, product showcases, and short-form creative content.

Fast, reliable workflows

Wan provides text-to-video (T2V), image-to-video (I2V), and reference-to-video (R2V) tools that let users create content quickly and reliably. This makes it suitable for workflows where consistency and repeatable results are important.