SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning
2026-08-17 • Sound
SoundComputer Vision and Pattern Recognition
AI summaryⓘ
The authors created SingDance, a new system that makes videos of people singing and dancing based on an image, text, and music. Unlike previous methods that only focus on dancing or assume the person is speaking, this system treats the visible person as either the singer or a listener, allowing more control over how they move and articulate vocals. They trained the system on different video types but never on videos that show singing and dancing together, yet it can still generate these videos at run-time. Their experiments showed that SingDance aligns movements well with music and lip-syncs effectively while using fewer resources than other speech-driven models.
video diffusionmusic-conditioned body motionvocal articulationspeech-driven modelslip synchronizationzero-shot learningsemantic roleaudio injectioncompositional generationmotion-beat alignment
Authors
Tao Feng, Xu Li, Xiangyang Luo, Ming Wen, Huadai Liu, Chen Zhang, Wei Xue
Abstract
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily on choreography, while speech-driven models generally assume that the visible subject produces the input voice, leaving this combined setting largely underexplored. We introduce SingDance, a unified video diffusion framework that formulates controllable vocal articulation as a semantic role: the visible subject is either the source, who produces the vocal signal, or the listener, who receives it from an off-screen performer. Hard-compact routing selects task-relevant speech, music, and role conditions, which are composed through frame-wise joint audio injection; source and listener retain the same speech pathway. Training uses asymmetric supervision: on-screen speaking and curated off-screen conversational-response videos establish role control, while instrumental and song-based dancing-only videos establish music-conditioned body motion. The target Song/Source configuration is never observed during training. At inference, assigning the source role to a song composes separately learned articulation and song-conditioned dance capabilities, enabling compositional zero-shot singing-and-dancing. Experiments demonstrate strong motion--beat alignment and visual fidelity, reliable paired switching of vocal articulation while preserving music-aligned body motion, and highly competitive lip synchronization with substantially fewer generation-time parameters than the strongest speech-driven baseline evaluated.