Speech enhancement models with PESQ loss score higher but sound worse
Perceptual Quality Loss or Loss of Perceptual Quality?
Machine Learning
Summary
Improving speech clarity using computer models is often measured by special scores like PESQ, which tries to predict how good speech sounds to humans. The authors found that training models to get higher PESQ scores doesn't always mean the speech sounds better to people. They ran tests where people listened to the outputs and preferred models trained without trying to improve PESQ. This shows that relying too much on one type of score can be misleading when making speech sound better.
What this means in practice
- •For speech software developers: Avoid using PESQ-optimized loss when training speech enhancement models to improve actual perceived sound quality.
- •For audio quality assurance teams: Use comprehensive listening tests alongside metric scores to evaluate speech enhancement models for real-world audio quality.
Authors
Danilo de Oliveira, Tal Peer, Maurício do V. M. da Costa, Timo Gerkmann
Abstract
Contemporary deep speech enhancement (SE) models are often trained with specific auxiliary terms in the loss function as a way to improve their performance in terms of perceptual metrics. Nevertheless, a higher score on a perceptual metric does not necessarily correlate with an improved listening experience. Through objective and subjective experiments, we assess the performance of SE models trained with two different types of auxiliary PESQ loss terms. The numerical evaluation on a suite of standard metrics suggests that, while models optimized for PESQ naturally obtain higher PESQ scores in the test set, for most other metrics the scores do not significantly change. In some cases, the PESQ loss even results in worse PESQ scores on mismatched data. A formal listening experiment reveals that the models without a PESQ loss were generally preferred over models that include it, across all settings. Finally, we analyze the relative importance of PESQ in the composite metrics CSIG, CBAK and COVL, and find that PESQ dominates all of them. Our study highlights the perils of over-reliance on PESQ and stresses the importance of a complete evaluation procedure for SE.