Audio visual models struggle with turn taking in noisy conversations
Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios
SoundComputation and Language
Summary
It's hard for computers to predict when someone will speak in noisy, crowded places like parties. The authors tested existing models that use sound and images, finding they don't work as well when background chatter or overlapping speech is present. They saw that retraining the models on noisy data helps, but some methods improve more than others depending on the training data and whether sound or visual information is used. This study shows it's important to make these models better at handling real-world noisy conversations.
What this means in practice
- •For conversational ai developers: Improve AI systems that predict when to speak by using audio-visual data adapted to noisy, overlapping speech in real-world environments.
- •For hearing aid engineers: Design hearing aids that can better anticipate speaker turn-taking amid background noise using improved multi-modal prediction models.
Authors
Long-Vu Hoang, Naomi Harte
Abstract
Current predictive turn-taking models (PTTMs) achieve strong performance on benchmarks with controlled acoustic conditions and clean audio signals. Their generalisation to conversations with overlapping speech and background interference remains underexplored. In this research, we evaluate audio-visual PTTMs trained with clean data on a challenging cocktail-party testbed derived from the AVCocktail dataset, and analyse their adaptation behaviour to this new domain. Experimental results show consistent performance degradation across audio and visual modalities under noisy conditions, with up to 38% relative drop in weighted F1. Fine-tuning on the new domain improves robustness, but gains vary across modalities and depend on the size of the available pre-training data. These findings provide insights into the different generalisation and adaptation capabilities of the audio and visual modalities, and indicate the need for robust modelling strategies to adapt to the complexities of human interactions in noise. All code and turn labels are made publicly available to facilitate further research.