Speech enhancer improves clarity and speaker identity without paired data
Unsupervised Speech Enhancement via Drifting
Sound
Summary
Cleaning up noisy speech recordings usually needs matching examples of clean and noisy sounds, which are hard to get. This paper shows a way to train speech enhancers using separate batches of clean and noisy audio without matching pairs. The authors fix problems with losing words and who is speaking by adding methods that connect the enhanced audio back to the original noisy input. Their method improves understanding and keeps voices more recognizable, all without needing labels or matched training data.
What this means in practice
- •For voice assistant developers: Improve voice assistants’ ability to understand speech in noisy environments without needing paired recordings for training.
- •For telecommunications engineers: Enhance call clarity and preserve speaker identity in noisy or reverberant calls without expensive data collection.
Authors
Diego Caviedes-Nozal, Liang Xu, Rasmus Kongsgaard Olsson, W. Bastiaan Kleijn
Abstract
This paper addresses unsupervised speech enhancement in the unpaired setting using drifting methods, where training relies on separate collections of degraded and clean audio without corresponding pairs. While recent drifting approaches enable unpaired training, they do so at a heavy cost: because the objective optimizes only a marginal prior over clean speech, the enhancer gradually loses the input's linguistic content and speaker identity. To fix this, we introduce input-conditioned drifting. We preserve the pull of the clean corpus while re-tethering the output to the degraded input via two mechanisms: an anchor encoder supplies the missing likelihood by pulling toward the input's features, and a key encoder conditions the prior by re-weighting retrieved frames. Neither requires labels or paired data. Using a training-free encoder selection criterion, Word Error Rate on VoiceBank-DEMAND falls to 10.1% (unprocessed: 11.7%), speaker similarity recovers from 0.490 to 0.879, and the recipe transfers in part to dereverberation on WSJ0-REVERB: content improves, rendering quality does not.