Articulatory models enable clear measurement of accent differences

Flexible and Interpretable Accent Distance Measurements

Artificial Intelligence

Summary

Measuring how different one person's accent is from another's can be tricky and depends on how you do it. The authors show that using models of how speech sounds are physically made in the mouth gives a way to compare accents that is both flexible and easy to understand. They also use a mathematical method called optimal transport to compare these speech features from any kind of recording. This provides a new way to study accents without needing special recordings or losing clarity about what the differences mean.

What this means in practice

Authors

Charles McGhee, Mark J. F. Gales, Kate M. Knill

Abstract

Determining the differences between two speakers' accents is a fundamental task in linguistics and speech technology research. The methodology used to measure these differences depends on the specific research area. A phonetics researcher may demonstrate accent variation by comparing vowel formants in paired recordings of individual words. These results will be interpretable, but the recordings will be time-consuming to collect and may not be representative of connected speech. Accented Text-to-Speech (TTS) research has pushed towards using accent embeddings derived from accent classification tasks. These embeddings can be produced from any speech recording, but are not readily interpretable. In this paper, we demonstrate that articulatory representations created through articulatory inversion can be used as an interpretable basis for accent comparison and that optimal transport provides a framework for accent comparison across arbitrary recording types.