Audio visual models mistakenly match speech to faces by position
Uncovering Ordinal-Matching Bias in Audio-Visual LLMs
Computer Vision and Pattern RecognitionSound
Summary
Audio-visual large language models often try to match spoken words to faces in videos, but the authors found they usually just pair speech to the face's position on screen instead of using real audio and visual cues. They call this mistake ordinal-matching bias. To fix this, the authors created a new training method that randomizes speaker positions and speaking order during fine-tuning, which helps the models learn to associate speech and faces correctly. This improvement makes the models better at understanding who’s speaking in real videos.
What this means in practice
- •For video conferencing developers: Improve speaker identification in multi-person video calls by reducing face-position bias in audio-visual models.
- •For video content indexing teams: Enhance automated tagging of who is speaking by training models to link speech and visible speakers accurately.
Authors
Jihoo Jung, Youngjoon Jang, Hyebin Cho, Suho Yoo, Joon Son Chung
Abstract
This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.