Speech recognition learns new words itself from unlabeled test audio

Learning New Words from Unlabeled Test Data in Automatic Speech Recognition

Computation and Language

Summary

Speech recognition systems usually struggle with new words they have never heard before. This paper shows how such systems can learn new words by listening to unlabeled speech data during testing, much like humans learn new words from context. The method uses a fixed speech-to-text model to generate spelling guesses and a language model to confirm the context of possible new words. Together, an adaptation module improves recognition of these new words without needing manual transcripts. Testing showed it reduces errors significantly on both regular and impaired speech datasets.

What this means in practice

  • For speech software developers: Improve speech systems' handling of new or rare words by enabling learning from unlabeled test speech audio without retraining.
  • For assistive technology teams: Enhance recognition of uncommon or personalized vocabulary in speech devices for users with speech impairments using unlabeled user data.

Authors

Mengqi Wang, Mark A. Hasegawa-Johnson, Haolong Zheng, Chang D. Yoo

Abstract

New words are invented every day. A human listener can learn a new word by hearing it clearly once and inferring its usage from sentence context. This paper proposes granting ASR a similar ability to learn the contextual representations and spellings of new words from unlabeled test data at test time. A frozen CTC acoustic model provides spellings, a frozen language model provides contextual evidence for out-of-vocabulary (OOV) word detection, and an adaptation module expands the vocabulary by learning the lexical token representations with distributions over CTC-generated candidates. The spelling model of each token is optimized by minimizing a Kullback-Leibler divergence (KLD) objective. We demonstrate that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and estimated spelling of the unknown word. Experiments show relative OOV character-error-rate reductions of up to 14.97% on LibriSpeech and 6.67% on dysarthric Speech Accessibility Project data for recurring OOV words, relative to the corresponding rescoring system.