Audio visual speech recognition struggles outside broadcast settings

AVSRBench: A Multi-Condition AVSR Benchmark

Computer Vision and Pattern RecognitionMultimedia

Summary

Speech recognition that uses both sound and lip movements works very well on TV broadcast speech, but the authors found it struggles when used in everyday situations like casual conversations or unusual speaking styles. They tested different systems in a variety of conditions and found that visual-only recognition fails quickly outside broadcast video, and even combined audio-visual systems mostly rely on sound when the speaker is not facing the camera. This shows current speech recognition technology does not generalize well to real-world video settings. The authors also offer a new dataset and tools to help researchers better test their systems in diverse conditions.

What this means in practice

Authors

Rishabh Jain, Naomi Harte

Abstract

While AVSR has achieved sub-1% word error rates on the standard LRS3 benchmark, its reliance on broadcast speech obscures whether this reflects true generalization or just domain adaptation. To investigate this gap, we evaluate three AVSR architectures across six conditions: controlled broadcast speech, fixed-grammar utterances, hyper-articulated Lombard speech, read speech from professional lipspeakers and non-professional speakers, and spontaneous multi-party video conversations. We find that visual-only performance deteriorates rapidly beyond broadcast domains, and audio-video fusion mainly benefits Lombard speech environments. Visual understanding degrades sharply at 90° profile views, with multimodal systems relying largely on acoustic fallback. Additionally, speaker articulation proves more critical than minor camera shifts, and LLM-based architectures suffer from poor out-of-domain generalization. Our work highlights a significant generalization gap in current AVSR research. To address this, we also introduce RoomReader-AV as a new benchmark for AVSR and release a unified data preprocessing pipeline to make comprehensive multi-condition evaluation accessible.