Decoupled Latent Flow Matching for Few-Step Joint Vocal-Accompaniment Separation
2026-08-31 • Sound
SoundMultimedia
AI summaryⓘ
The authors developed a method to separate singing voices from musical accompaniment using a special technique called latent flow matching. They first compress the music into a smaller, simpler form using a variational autoencoder, then use a model to separate the vocals and instruments in this compressed space. To make the process faster, they added a step inspired by another method (Flow2GAN) that helps produce results with fewer steps. Their experiments showed this approach improves the quality of separated sounds while being more efficient.
generative modelinglatent flow matchingvariational autoencoder (VAE)vocal-accompaniment separationFlow2GANsampling costsemantic separationacoustic velocity predictionlatent adversarial post-training
Authors
Lishi Zuo, Youzhi Tu, Lu Yi, Zezhong Jin, Chongxin Gan, Man-Wai Mak, KongAik Lee
Abstract
Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.