Papers for

speech software developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Speech enhancement models with PESQ loss score higher but sound worse

Perceptual Quality Loss or Loss of Perceptual Quality?

Abstract: Contemporary deep speech enhancement (SE) models are often trained with specific auxiliary terms in the loss function as a way to improve their performance in terms of perceptual metrics. Nevertheless, a higher score on a perceptual metric does not necessarily correlate with an improved listening experience. Through objective and subjective experiments, we assess the performance of SE models trained with two different types of auxiliary PESQ loss terms. The numerical evaluation on a suite of standard metrics suggests that, while models optimized for PESQ naturally obtain higher PESQ scores in the test set, for most other metrics the scores do not significantly change. In some cases, the PESQ loss even results in worse PESQ scores on mismatched data. A formal listening experiment reveals that the models without a PESQ loss were generally preferred over models that include it, across all settings. Finally, we analyze the relative importance of PESQ in the composite metrics CSIG, CBAK and COVL, and find that PESQ dominates all of them. Our study highlights the perils of over-reliance on PESQ and stresses the importance of a complete evaluation procedure for SE.

Mon 28 SeptMachine Learning
The gist
Improving speech clarity using computer models is often measured by special scores like PESQ, which tries to predict how good speech sounds to humans. The authors found that training models to get higher PESQ scores doesn't always mean the speech sounds better to people. They ran tests where people listened to the outputs and preferred models trained without trying to improve PESQ. This shows that relying too much on one type of score can be misleading when making speech sound better.
Open → 2609.35054v1

Speech recognition learns new words itself from unlabeled test audio

Learning New Words from Unlabeled Test Data in Automatic Speech Recognition

Abstract: New words are invented every day. A human listener can learn a new word by hearing it clearly once and inferring its usage from sentence context. This paper proposes granting ASR a similar ability to learn the contextual representations and spellings of new words from unlabeled test data at test time. A frozen CTC acoustic model provides spellings, a frozen language model provides contextual evidence for out-of-vocabulary (OOV) word detection, and an adaptation module expands the vocabulary by learning the lexical token representations with distributions over CTC-generated candidates. The spelling model of each token is optimized by minimizing a Kullback-Leibler divergence (KLD) objective. We demonstrate that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and estimated spelling of the unknown word. Experiments show relative OOV character-error-rate reductions of up to 14.97% on LibriSpeech and 6.67% on dysarthric Speech Accessibility Project data for recurring OOV words, relative to the corresponding rescoring system.

Thu 24 SeptComputation and Language
The gist
Speech recognition systems usually struggle with new words they have never heard before. This paper shows how such systems can learn new words by listening to unlabeled speech data during testing, much like humans learn new words from context. The method uses a fixed speech-to-text model to generate spelling guesses and a language model to confirm the context of possible new words. Together, an adaptation module improves recognition of these new words without needing manual transcripts. Testing showed it reduces errors significantly on both regular and impaired speech datasets.
Open → 2609.28877v1