Challenges improve speech extraction in multi-speaker conversations

Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

SoundComputation and Language

Summary

In noisy conversations with many people, pulling out one person’s voice clearly is hard, especially when speakers talk less and their reference samples differ from what they say in the conversation. The authors found that current methods trained on balanced, simulated data don’t work as well in real-life talks. They created a new training technique to handle long silences better, improving the clarity and quality of the extracted speech. They also studied how mismatched reference samples affect performance.

What this means in practice

Authors

Robert Sutherland, Stefan Goetze, Jon Barker

Abstract

Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.