Alibaba's Wan Video Turns Any Photo Into a Music-Synced Dance Clip

Alibaba's Wan Video adds music-driven dance generation: upload a character image, pick a style, and get a beat-synced video in seconds

ByWanWan
·
·
AuthorWan
Read2 min
  • New feature: Wan Video's Music to Dance generates beat-synchronized dance videos from a character image and a song.
  • Five styles available: Street, Tap, Latin, K-Pop, and Chinese Classical — more styles expected as the feature matures.
  • Live now: Available on create.wan.video today; mobile app version coming soon.
  • Pricing: Free tier available; Pro plan starts at $5/month (billed yearly) with 300 credits (~60 videos); commercial rights require a paid plan.
  • Key differentiator: Music-first workflow generates motion that responds to the uploaded audio, rather than applying a fixed choreography template.
  • Best for: Faceless creators, brand mascot animation, fan edits, and social content synced to specific tracks — no animation skills required.

Wan Video, Alibaba's AI creative platform, just shipped a feature that skips the hardest part of dance content creation: the dancing. The new Music to Dance tool lets you upload a character image, drop in a song, and get back a fully choreographed video where your character moves in sync with the music. No motion capture, no keyframing, no choreographer on retainer.

What it actually does

The workflow is three steps: upload a character image, upload or select a song, and pick a dance style. Wan is an AI creative platform that aims to lower the barrier to creative work, offering features like text-to-image, image-to-video, and image editing. Music to Dance is its most opinionated feature yet , it takes the audio input and generates motion that is explicitly synchronized to the rhythm of the track.

Five dance styles are currently supported:

  • Street
  • Tap
  • Latin
  • K-Pop
  • Chinese Classical

The character image can be a photo, an illustration, or an AI-generated avatar. You can use portraits, half-body, or full-body shots, and the model works with photos, illustrations, cartoons, and more. The key constraint is that the image should feature a single, clearly visible character , the cleaner the input, the better the output.

The synchronization problem it solves

Getting motion to match music is genuinely hard. The standard approach in research involves extracting beat timestamps from the audio signal, then aligning movement keypoints to those timestamps. To align the temporal rhythm of a video with a target music track, systems identify and match rhythmic structures in both modalities , extracting music beats from the audio and salient motion points from the video, then computing a one-to-one temporal alignment between them. Doing this at generation time, rather than as a post-processing step, is what makes the output feel natural rather than mechanically edited.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves