Articulatory features improve pathological speech intelligibility assessment
ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment
Sound
Summary
Measuring how clear speech is from people with speech difficulties is important but challenging because tools need to be both accurate and easy to understand. The authors replaced complex audio features with predicted mouth movement patterns in an existing speech clarity measure. This new approach, ART-NAD, matches older methods in accuracy but adds the benefit of showing which specific mouth movements cause problems. This extra detail helps specialists see exactly where speech breaks down.
What this means in practice
- •For speech therapy clinics: Assess speech intelligibility using interpretable metrics that highlight specific articulatory issues in pathological speakers.
- •For speech technology developers: Incorporate articulatory inversion models to build clearer feedback tools that visually map speech articulation problems for clinical use.
Authors
Bence Mark Halpern, Thomas Tienkamp, Defne Abur, Tomoki Toda
Abstract
Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-supervised features that are hard to interpret, providing only frame-level explanations. We propose ART-NAD, a reference-audio intelligibility metric that replaces the \texttt{wav2vec2} features of NAD with vocal-tract constriction variables (tract variables, TVs) predicted from audio by a speaker-independent acoustic-to-articulatory inversion model trained on the same \texttt{wav2vec2} features. ART-NAD is computed as the multivariate Dynamic Time Warping distance between the nine-channel quasi-TV trajectories of the test and one or more references. Across 20 reference-audio protocols spanning six pathological-speech datasets and five languages, ART-NAD with silence trimming (ART-NAD-FA) reaches the same average speaker-level Pearson correlation as NAD-FA (both $r=0.71$) on the same self-supervised backbone, with no significant per-protocol difference (Wilcoxon $p=0.18$), and is the strongest reference-audio metric on 6 of the 20 protocols. Beside the score itself, each TV channel visualizes which constriction deviates from the reference over time, providing interpretable information as to where articulation breaks down.