Visual and text guided method removes sounds linked to video objects
TV-AudioRemover: Joint Text-Visual Guided Sound Removal with Multi-Task Hard-Mixture Curriculum
MultimediaComputer Vision and Pattern RecognitionSound
Summary
When people remove objects from a video, the sounds those objects made often remain, making the video feel strange. The authors created a method that uses both the edited video and text instructions to remove the sounds linked to the removed objects. They trained their model on a large set of video and audio clips with aligned objects and sounds. Their approach improves how well the unwanted sounds are canceled while keeping the rest intact. They also created a new benchmark to measure how well these methods work.
What this means in practice
- •For video editors: Remove sounds associated with objects deleted from videos to keep audio consistent with visuals.
- •For audio production teams: Clean audio tracks by suppressing sounds connected to visual elements removed in post-processing.
- •For movie special effects teams: Enhance special effects by jointly editing video and removing related environmental sounds for seamless scenes.$Commercial implications: Enables production companies to offer advanced post-production tools that improve audio-visual consistency in films.
Authors
Xinyue Guo, Jianxuan Yang, Daiguo Zhou, Jiagao Hu, Yuxuan Chen, Fei Wang, Jian Luan
Abstract
Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and therefore rely on limited single-modal control, which is less effective than multimodal guidance that provides stronger semantic grounding and temporal synchronization cues. In this paper, we present Text-Visual Guided Sound Removal (TV-AudioRemover), a target sound removal framework that leverages the visually edited video together with a natural-language instruction to suppress the sound associated with the removed visual object from the original audio mixture. To acquire high-quality training data, we devise a pipeline to construct a million-scale dataset of single-object audio-visual aligned samples, from which we synthesize mixture-target pairs customized for model training. To effectively leverage visual context and follow instruction intent, we augment the model architecture with task tokens, generalizable instruction modeling, and modality-specific global guidance. We further adopt multi-task training to strengthen task-role comprehension, and employ a hard-mixture curriculum that leverages semantically similar acoustic mixtures during fine-tuning to enhance fine-grained source discrimination. To support evaluation, we present AV-Remove-Bench, a comprehensive audio-visual object removal benchmark, along with dedicated objective metrics and an MLLM-based evaluation protocol. Experiments demonstrate that our method achieves state-of-the-art performance on both subjective and objective metrics. Project page: https://yjx-research.github.io/TV-AudioRemover/.