Linguistic Distance Segregates Latent Representations in Automatic Speech Recognition Systems

2026-08-31Computation and Language

Computation and LanguageMachine Learning
AI summary

The authors studied how well speech-to-text systems understand English speakers whose first languages are very different from English. They found that the more different a speaker's first language is, the more mistakes the system makes. This pattern was consistent across different datasets and models and was statistically significant. They also discovered that deeper layers in the models show signs of grouping speech based on the speaker’s first language.

automatic speech recognitionfirst language (L1)ASR error ratelanguage distanceTweedie mixed-effects modellatent spaceacoustic layersstatistical significancedataset variationspeech processing
Authors
Ting-Hui Cheng, Line Katrine Harder Clemmensen, Sneha Das
Abstract
While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across different speaker populations. One such disparity is for speakers whose first languages (L1) are from families distant from English. This paper investigates the relationship between first language background and English ASR performance. Through empirical analysis, we observe that the correlation between speakers' L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models. This association is statistically significant in a follow-up analysis accounting for dataset-level variation in Tweedie mixed-effects models ($p<0.001$ across evaluated models). In addition, analysis of the latent space reveals a L1-based spatial segregation across deeper acoustic layers in the majority of evaluated architectures