Benchmark measures code switching errors impact on enterprise voice systems
CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
SoundComputation and LanguageMachine Learning
Summary
Automatic speech recognition systems struggle when people switch languages mid-sentence. The authors created a test set and evaluation tools to see how these mistakes affect business voice assistants. They checked current speech systems on five language pairs and studied what extra errors code-switching causes. Their work helps make voice agents better at understanding mixed-language speech in workplace settings.
What this means in practice
- •For enterprise voice technology teams: Evaluate how well speech recognition handles mixed languages and how errors affect voice assistant tasks in business environments.
- •For multilingual customer support teams: Assess voice agent performance on code-switched speech to improve customer interactions across languages in call centers.
Authors
Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin, Fanny Riols
Abstract
Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how code-switching transcription errors propagate to downstream voice agent task failures. In this work, we propose (1) a CS-ASR synthetic benchmark and multidimensional evaluation framework tailored to enterprise domains, (2) systematic evaluation of frontier ASR systems across 5 language pairs, (3) diagnostic analysis of the additional transcription errors that code-switching introduces across language pairs and models. We release COSE-E to support enterprise-focused CSASR evaluation for multilingual voice agents in enterprise deployment.