Benchmark measures code switching errors impact on enterprise voice systems

CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings

SoundComputation and LanguageMachine Learning

Summary

Automatic speech recognition systems struggle when people switch languages mid-sentence. The authors created a test set and evaluation tools to see how these mistakes affect business voice assistants. They checked current speech systems on five language pairs and studied what extra errors code-switching causes. Their work helps make voice agents better at understanding mixed-language speech in workplace settings.

What this means in practice

Authors

Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin, Fanny Riols

Abstract

Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how code-switching transcription errors propagate to downstream voice agent task failures. In this work, we propose (1) a CS-ASR synthetic benchmark and multidimensional evaluation framework tailored to enterprise domains, (2) systematic evaluation of frontier ASR systems across 5 language pairs, (3) diagnostic analysis of the additional transcription errors that code-switching introduces across language pairs and models. We release COSE-E to support enterprise-focused CSASR evaluation for multilingual voice agents in enterprise deployment.