Audio models reveal hidden errors in code switched English Yoruba speech
Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech
Summary
Speech recognition systems usually perform well on single languages, but this paper shows that when people mix English and Yoruba in the same sentence, the usual error scores don’t tell the full story. The authors tested eleven modern speech systems on mixed English-Yoruba speech and found that while overall error rates look good, these systems struggle to recognize Yoruba words correctly and make many mistakes at language switching points. Some systems even add extra words or change the meaning depending on how the test questions are framed. The authors provide tools to better measure these mistakes in mixed language speech.
What this means in practice
- •For speech technology developers: Develop more accurate speech recognition tools by using switch-aware metrics to detect errors specifically at points where languages change in mixed speech.
- •For voice assistant engineers: Improve mixed language understanding in voice interfaces by using refined evaluation methods that identify failures in recognizing low-resource language words and language switches.