Papers for

media indexing teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Active speaker detection improves with joint audio and face modeling

ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion

Abstract: Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.

Tue 22 SeptMultimediaComputer Vision and Pattern RecognitionSound
The gist
Active speaker detection identifies who is speaking by linking what you hear with the faces you see. Existing methods struggle when video or audio is unclear or incomplete. The authors introduce ROAM-ASD, which looks at full faces, mouth details, and sounds together using a system that pays attention to all these inputs at once. This approach is better at detecting speakers even when some information is missing, working well across several test collections.
Open → 2609.26648v1