WhisperVC-AV improves noisy whisper to normal voice conversion
WhisperVC-AV: Audio-Visual Content Restoration for Noise-Robust Whisper-to-Normal Voice Conversion
Sound
Summary
Turning whispered speech into normal talking voice is hard when there’s background noise because the sounds are harder to understand. The authors created WhisperVC-AV, which uses both the sound and lip movements to better guess what was said. This method helps computers recognize whispered words more accurately, even with lots of noise, without changing the main conversion process. It also keeps the speaker’s voice sounding natural.
What this means in practice
- •For speech technology developers: Improve voice conversion systems to handle noisy whispered inputs by integrating lip movement cues without altering core conversion modules.
- •For hearing aid designers: Enhance whispered speech clarity in hearing aids by combining audio and visual lip information to reduce errors caused by background noise.
Authors
Ziyue Yin, Dong Liu, Ming Li
Abstract
Environmental noise obscures the content cues needed for whisper-to-normal conversion. We propose WhisperVC-AV, which restores acoustic content features while leaving the original WhisperVC conversion module unchanged. Its context-guided restoration module combines temporal acoustic context with synchronized lip features through attention and gated residual correction. Experiments on AISHELL6-Whisper show lower character error rates (CERs) than WhisperVC across three ASR systems, on clean speech and under all six signal-to-noise ratio (SNR) conditions with MUSAN noise. The largest gains occur at 0 dB SNR, where Qwen3-ASR CER falls from 36.29% to 28.88%. WhisperVC-AV also improves predicted speech quality while maintaining speaker similarity. The CER gains extend to unseen background noise without retraining, while visual controls support the use of utterance-specific lip cues. Audio examples are available on our demo page.