Lightweight model separates singing voices using audio and video cues
MambaVoice: Lightweight Audiovisual Singing Voice Separation Via A Hybrid Mamba-Transformer Model
Sound
Summary
Separating a singer’s voice from music videos is hard, especially when many people sing or lots of instruments play. The authors created MambaVoice, a compact model that uses both sound and face movements to pick out the target singer’s voice. It combines audio and video information in a clever way to focus on the right voice, and uses advanced methods to understand long stretches of music efficiently. Tests show MambaVoice works well compared to bigger models but with fewer resources.
What this means in practice
- •For music software developers: Integrate MambaVoice to separate individual singing voices in music videos with low computational resources and high accuracy.
- •For video editors: Use audiovisual cues to isolate or mute specific singers in multi-vocal music videos for clearer sound editing.
Authors
Adithi Shankar, Gopika Krishnan, Gloria Haro, Xavier Serra, Martín Rocamora
Abstract
Isolating a target singing voice from a music video remains challenging, particularly in the presence of multiple vocalists and dense instrumental accompaniment. We propose MambaVoice, a lightweight audiovisual framework that leverages a hybrid Mamba--Transformer architecture for targeted singing voice separation. The model jointly encodes audio and visual streams using an attention-based band-split audio encoder and a spatio-temporal graph convolutional network (ST-GCN) for facial motion features. These modalities are fused through a multiplicative gating mechanism, enabling visual cues to selectively modulate audio representations. The fused features are processed by a hybrid backbone that combines Transformer self-attention with Selective State Space Models (SSMs), achieving efficient long-range temporal modeling with linear complexity. We evaluated MambaVoice on the Acappella and URSing datasets under challenging conditions, including mixtures with interfering singers. At 16.2 million parameters, the model demonstrates comparable performance, achieving 14.18 dB SDR on Acappella and strong cross-dataset performance on URSing, comparable to larger models at a fraction of the parameter count. These findings highlight the effectiveness of hybrid SSM--attention architectures for scalable, efficient audiovisual source separation, suggesting they are well-suited as lightweight components within larger pipelines. We conduct a perceptual study that further supports our improvements in objective metrics. We provide our implementation online.