Papers for

video conferencing developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Audio visual models mistakenly match speech to faces by position

Uncovering Ordinal-Matching Bias in Audio-Visual LLMs

Abstract: This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.

Mon 28 SeptComputer Vision and Pattern RecognitionSound
The gist
Audio-visual large language models often try to match spoken words to faces in videos, but the authors found they usually just pair speech to the face's position on screen instead of using real audio and visual cues. They call this mistake ordinal-matching bias. To fix this, the authors created a new training method that randomizes speaker positions and speaking order during fine-tuning, which helps the models learn to associate speech and faces correctly. This improvement makes the models better at understanding who’s speaking in real videos.
Open → 2609.34223v1

Audio-visual method enhances target speaker voice with low delay

Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction

Abstract: Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.

Thu 24 SeptComputer Vision and Pattern RecognitionSound
The gist
When there are multiple people talking, it can be hard to focus on just one person's voice clearly. The authors designed a system that uses both sound and video (like watching mouth movements) to cleanly extract one speaker's voice from a mix of voices in real time. Their method improves the quality of the chosen speaker's voice and reduces confusion from other voices, doing this quickly with minimal delay. They tested it on real meeting recordings and listening tests, showing better results than previous approaches.
Open → 2609.30631v1

Active speaker detection improves with joint audio and face modeling

ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion

Abstract: Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.

Tue 22 SeptMultimediaComputer Vision and Pattern RecognitionSound
The gist
Active speaker detection identifies who is speaking by linking what you hear with the faces you see. Existing methods struggle when video or audio is unclear or incomplete. The authors introduce ROAM-ASD, which looks at full faces, mouth details, and sounds together using a system that pays attention to all these inputs at once. This approach is better at detecting speakers even when some information is missing, working well across several test collections.
Open → 2609.26648v1

Audio visual speech recognition struggles outside broadcast settings

AVSRBench: A Multi-Condition AVSR Benchmark

Abstract: While AVSR has achieved sub-1% word error rates on the standard LRS3 benchmark, its reliance on broadcast speech obscures whether this reflects true generalization or just domain adaptation. To investigate this gap, we evaluate three AVSR architectures across six conditions: controlled broadcast speech, fixed-grammar utterances, hyper-articulated Lombard speech, read speech from professional lipspeakers and non-professional speakers, and spontaneous multi-party video conversations. We find that visual-only performance deteriorates rapidly beyond broadcast domains, and audio-video fusion mainly benefits Lombard speech environments. Visual understanding degrades sharply at 90° profile views, with multimodal systems relying largely on acoustic fallback. Additionally, speaker articulation proves more critical than minor camera shifts, and LLM-based architectures suffer from poor out-of-domain generalization. Our work highlights a significant generalization gap in current AVSR research. To address this, we also introduce RoomReader-AV as a new benchmark for AVSR and release a unified data preprocessing pipeline to make comprehensive multi-condition evaluation accessible.

Wed 9 SeptComputer Vision and Pattern RecognitionMultimedia
The gist
Speech recognition that uses both sound and lip movements works very well on TV broadcast speech, but the authors found it struggles when used in everyday situations like casual conversations or unusual speaking styles. They tested different systems in a variety of conditions and found that visual-only recognition fails quickly outside broadcast video, and even combined audio-visual systems mostly rely on sound when the speaker is not facing the camera. This shows current speech recognition technology does not generalize well to real-world video settings. The authors also offer a new dataset and tools to help researchers better test their systems in diverse conditions.
Open → 2609.10366v1