Papers for

hearing aid engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multimodal cues improve target voice extraction in noisy settings

Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities

Abstract: Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems conditioned on enrollment speech have progressed substantially, enrollment speech as a cue has inherent limitations. Its reliability degrades when the target and interfering speakers have similar voice characteristics, when intra-speaker variability (e.g. changes in emotion or speaking style) creates a mismatch between the enrollment and target speech, or when the enrollment itself is contaminated by noise or competing speakers. This review surveys deep-learning-based TSE from the perspective of auxiliary target cues drawn from multiple modalities. We organize existing methods according to five types of information used to isolate the target speaker: audio enrollment, visual, spatial, textual/semantic, and neural cues. We also trace the evolution from discriminative estimators to variational, diffusion, flow, codec, and foundation-model-based systems and summarize representative datasets and evaluation metrics. We review the benefits and limitations of different cues and discuss challenges involving synchronization, missing or unreliable observations, data scarcity, privacy, computational cost, and real-time operation. Finally, we summarize future directions concerning adaptive cue fusion, instruction-driven extraction, realistic evaluation, and trustworthy deployment. By jointly reviewing cue design, model architecture, training objectives, datasets, and evaluation metrics, this article provides an overview of the current landscape and open problems in multimodal TSE.

Mon 28 SeptSoundMultimedia
The gist
Extracting a single person's voice from a noisy room with many people talking is hard, especially if voices sound alike or the example voice recording is unclear. The authors review research that uses different kinds of information, like visual cues, location, text, and brain signals, to better isolate a speaker’s voice. They explain how different approaches and technologies have developed over time and highlight the challenges when some cues are missing or unreliable. Their work helps show what methods work best and where more progress is needed for practical, trustworthy voice separation.
Open → 2609.35613v1

Audio visual models struggle with turn taking in noisy conversations

Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios

Abstract: Current predictive turn-taking models (PTTMs) achieve strong performance on benchmarks with controlled acoustic conditions and clean audio signals. Their generalisation to conversations with overlapping speech and background interference remains underexplored. In this research, we evaluate audio-visual PTTMs trained with clean data on a challenging cocktail-party testbed derived from the AVCocktail dataset, and analyse their adaptation behaviour to this new domain. Experimental results show consistent performance degradation across audio and visual modalities under noisy conditions, with up to 38% relative drop in weighted F1. Fine-tuning on the new domain improves robustness, but gains vary across modalities and depend on the size of the available pre-training data. These findings provide insights into the different generalisation and adaptation capabilities of the audio and visual modalities, and indicate the need for robust modelling strategies to adapt to the complexities of human interactions in noise. All code and turn labels are made publicly available to facilitate further research.

Tue 15 SeptSoundComputation and Language
The gist
It's hard for computers to predict when someone will speak in noisy, crowded places like parties. The authors tested existing models that use sound and images, finding they don't work as well when background chatter or overlapping speech is present. They saw that retraining the models on noisy data helps, but some methods improve more than others depending on the training data and whether sound or visual information is used. This study shows it's important to make these models better at handling real-world noisy conversations.
Open → 2609.17056v1