Ensemble-Based Perceptual Audio Quality Assessment with Confidence Intervals

Sound

Summary

The gist is being written…

Authors

Pablo M. Delgado, Andreas Brendel, Konstantin Schmidt, Jürgen Herre

Abstract

Objective audio quality metrics typically provide point estimates, whereas listening tests yield score distributions from which mean opinion scores (MOS), confidence intervals (CIs), and significance decisions are derived. We propose a lightweight intrusive metric that combines a PEAQ-style perceptual front-end (ITU-R BS.1387) with a bagging ensemble of regressors. Calibration with subjective data aligns ensemble outputs with listener scores. The resulting item-dependent score distributions enable uncertainty assessment, panel-size-matched CIs, and identification of less conclusive predictions. Calibration improves agreement with subjective distributions and CI coverage across all evaluated datasets while preserving MOS accuracy. Using only 11 fixed PEAQ features, low-capacity regressors, and public training data, the method performs comparably to more data-intensive end-to-end approaches. Its output can support uncertainty-aware assessment and target listening tests towards uncertain conditions. The distributions also enable approximate pairwise comparisons, but not yet reliable significance inference.