Audio-visual method enhances target speaker voice with low delay

Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction

Computer Vision and Pattern RecognitionSound

Summary

When there are multiple people talking, it can be hard to focus on just one person's voice clearly. The authors designed a system that uses both sound and video (like watching mouth movements) to cleanly extract one speaker's voice from a mix of voices in real time. Their method improves the quality of the chosen speaker's voice and reduces confusion from other voices, doing this quickly with minimal delay. They tested it on real meeting recordings and listening tests, showing better results than previous approaches.

What this means in practice

  • For video conferencing developers: Enable real-time extraction of a participant’s voice from multi-speaker meetings to improve audio clarity with minimal delay.$Commercial implications: This paper enables higher-quality live audio enhancement products for video conference platforms, improving user experience in noisy group calls.
  • For hearing aid designers: Integrate audio-visual cues to better isolate and enhance a specific speaker’s voice in environments with competing talkers.

Authors

Rayhan Rashed, Senja Filipi, Ross Cutler

Abstract

Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.