What Changed
Wan-AI has released Wan-Dancer-14B, a novel image-to-video generation model specifically designed for music-to-dance synthesis. This model distinguishes itself by employing a hierarchical framework that addresses the challenge of generating long-duration, coherent dance videos. Unlike single-stage generation approaches, Wan-Dancer-14B separates the process into two distinct phases: global keyframe planning and local temporal refinement. This allows the model to maintain both overall structural consistency and fine-grained rhythmic accuracy over extended video sequences.
The release includes the model weights and inference code, making it accessible for developers to experiment with and integrate into their projects. The model supports various dance genres, including Chinese Classical Dance, K-Pop Dance, Street Dance, Latin Dance, and Tap Dance, by utilizing specific prompt files for each style.
Technical Details
Wan-Dancer-14B operates on a two-stage hierarchical framework. The core idea is to first establish the global structure and then refine the local temporal details, leveraging the full-track musical context to ensure long-range coherence in the generated dance.







