Iterative audio separation improves sound clarity using multi input output models
Iterative Audio Separation with Mixture Consistency via MIMO Model Extension
SoundMachine Learning
Summary
Separating different sounds in an audio mix can be tricky, especially when you want to keep the natural quality and timing of the sounds. This paper shows a new way to improve this process by using models that take multiple audio inputs and outputs at once, instead of just one. By doing this, the researchers were able to separate sounds more accurately through repeated steps without losing quality. They tested their approach on top existing methods and found it made the results noticeably better.
audio separationmixture consistencymulti-input multi-output (MIMO)source separationiterative predictionspeech enhancementdiffusion modelsgenerative modelsdiscriminatorsphase information
Authors
Yukara Ikemiya, WeiHsiang Liao, Yuki Mitsufuji
Abstract
This paper proposes a general framework for stable and effective iterative audio separation with mixture consistency by extending source separation models to a multi-input multi-output (MIMO) configuration. In the field of audio separation, mixture consistency is an essential property for many applications that require accurate phase and timbral information of target sources. While iterative approaches such as diffusion models achieve perceptually superior results in speech enhancement or user-guided target source separation tasks, most existing methods focus on single-step separation with a single-input single-output (SISO) or single-input multi-output (SIMO) configuration through architectural improvements, since mixture-consistent audio separation is generally regarded as a regression problem that admits a unique solution. By extending these architectures to a MIMO configuration, we introduce iterative prediction without compromising the architectural advantages or the characteristics of mixture consistency. We conduct a comprehensive ablation study of combining the framework with discriminators and extending it to a generative model. Experimental results demonstrate significant performance improvements when applying the proposed framework to state-of-the-art separation models.