Active speaker detection improves with joint audio and face modeling
ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion
MultimediaComputer Vision and Pattern RecognitionSound
Summary
Active speaker detection identifies who is speaking by linking what you hear with the faces you see. Existing methods struggle when video or audio is unclear or incomplete. The authors introduce ROAM-ASD, which looks at full faces, mouth details, and sounds together using a system that pays attention to all these inputs at once. This approach is better at detecting speakers even when some information is missing, working well across several test collections.
What this means in practice
- •For video conferencing developers: Enhance speaker tracking in calls by accurately linking voices to faces despite poor video or audio conditions.$Commercial implications: This enables more reliable speaker recognition features in conferencing products, improving user experience and engagement.
- •For media indexing teams: Automatically tag speakers in videos across diverse datasets, even when some audio or visual data is missing.
Authors
Pu Wang, Yujun Wang, Hugo Van hamme
Abstract
Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.