Binaural audio improves spotting sound sources in videos
Binaural Audio-Visual Instance Segmentation
Computer Vision and Pattern RecognitionSound
Summary
It's hard for computers to tell apart objects that look very similar but make different sounds just by watching videos with one microphone. The authors help computers by using two microphones, like how human ears work, to figure out where sounds are coming from more precisely. They created new tests and a method that uses these two microphones to better find and separate the exact sound-making objects in videos. Their method works better than older ones, especially when objects look alike but sound different.
What this means in practice
- •For video editing teams: Segment and isolate specific sound-making objects in videos with interactive editing tools using binaural audio.$Commercial implications: Enables advanced video editing software for enhanced audio-visual object separation by integrating binaural spatial cues.
- •For robotics developers: Improve robots’ ability to localize and distinguish multiple sound sources visually and auditorily using binaural cues.
Authors
Saijun Wang, Guanfeng Tang, Hongbo Zhao, Zhicheng Lei, Yutong Zhang, Wei Ye, Rui Fan
Abstract
Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually similar instances of the same semantic class. In contrast, humans naturally exploit binaural hearing, where interaural differences and direction-dependent acoustic filtering introduced by the head and pinnae provide physically grounded spatial cues for accurate sound source localization. Motivated by this observation, we introduce binaural audio-visual instance segmentation (BiAVIS), a new task that leverages synchronized binaural audio and video frames to segment sounding instances. To advance research on this task, we establish two benchmarks by manually annotating an existing binaural audio-visual dataset and collecting a new real-world dataset, BiAVIS-Bench, in more challenging and diverse scenarios. We further propose a BiAVIS model, which leverages an audio-only sound source localization network to learn spatial and semantic priors for sounding instances from binaural audio. A query-level audio-visual fusion strategy is subsequently introduced to inject these informative priors into the instance segmentation decoder. Extensive experiments conducted on the two proposed benchmarks demonstrate the superior performance of the BiAVIS model over previous monaural AVS methods, especially in resolving instance-level intra-class ambiguity. On the more challenging BiAVIS-Bench, the proposed BiAVIS model outperforms the best-performing monaural baselines by 17.22\% in mAP and 7.71\% in FSLA, respectively.