SSE model enhances and remixes audio using video and text guidance
Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing
SoundArtificial Intelligence
Summary
Mixing audio in videos is hard when unwanted sounds or echo get in the way. The authors created a new system called SSE that uses the video and text descriptions to find, separate, and improve sounds in a video’s audio track. This lets people remove unwanted noise, balance sound levels, and reduce echoes with simple instructions. They also made a new dataset to help train and test their system and showed it works better than existing methods.
What this means in practice
- •For video editors: Adjust audio tracks in video clips by selectively removing unwanted sounds and rebalancing audio using both video scenes and text instructions.
- •For game audio designers: Create cleaner and more controllable in-game sound environments by separating and enhancing audio elements guided by visual and textual context.
Authors
Ilpo Viertola, Giulio Cengarle, Gouthaman KV, Daniel Arteaga, Lie Lu
Abstract
We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site