Papers for

multilingual customer support teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Benchmark measures code switching errors impact on enterprise voice systems

CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings

Abstract: Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how code-switching transcription errors propagate to downstream voice agent task failures. In this work, we propose (1) a CS-ASR synthetic benchmark and multidimensional evaluation framework tailored to enterprise domains, (2) systematic evaluation of frontier ASR systems across 5 language pairs, (3) diagnostic analysis of the additional transcription errors that code-switching introduces across language pairs and models. We release COSE-E to support enterprise-focused CSASR evaluation for multilingual voice agents in enterprise deployment.

Mon 28 SeptSoundComputation and LanguageMachine Learning
The gist
Automatic speech recognition systems struggle when people switch languages mid-sentence. The authors created a test set and evaluation tools to see how these mistakes affect business voice assistants. They checked current speech systems on five language pairs and studied what extra errors code-switching causes. Their work helps make voice agents better at understanding mixed-language speech in workplace settings.
Open → 2609.35645v1

Arabic speech recognition benchmark spans 17 dialects and cultures

Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition

Abstract: Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.

Mon 28 SeptComputation and LanguageArtificial Intelligence
The gist
Most Arabic speech recognition technology works well only for the formal written language, Modern Standard Arabic, not the many spoken dialects. The authors created ALMIEYAR, a new test set with recordings from 17 different Arabic dialects, based on native speakers describing local cultural images across everyday topics. They tested 12 advanced speech recognition systems on it and found that all still make a lot of mistakes, especially because dialects vary widely. This resource helps measure how well speech recognition systems understand different Arabic dialects and highlights where improvements are needed.
Open → 2609.35564v1

Large language models improve real time speech translation with less delay

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

Abstract: Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.

Thu 24 SeptComputation and Language
The gist
Translating spoken language in real time is hard because it requires quickly understanding and converting speech while keeping the speaker’s voice. The authors developed a new way to train language models that respect the timing of speech and keep the speaker’s voice identity. This method uses less training data but still delivers better translation quality and faster responses in multiple languages. Their system balances accuracy and speed better than previous methods, helping conversations flow more naturally across languages.
Open → 2609.30416v1

Joint training improves speech recognition and translation accuracy

Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation

Abstract: In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimization (GRPO). We score both transcripts and translations, with translation conditioned on model-generated transcripts, and compare three token advantage strategies. Using Qwen2.5-Omni-3B across four languages, we evaluate CoT against direct speech translation (Direct ST) under SFT and GRPO, training on CoVoST 2 and testing on CoVoST 2 and FLEURS. CoT GRPO outperforms Direct ST GRPO by 1.77 and 0.83 average BLEU points on CoVoST 2 and FLEURS. Compared to CoT SFT, GRPO boosts BLEU by 0.82 and 0.67 points and reduces word error rate (WER) by 8.8% and 7.2% relatively. These results highlight reinforcement fine-tuning as an effective method to mitigate the training-inference mismatch, jointly improving recognition and translation.

Tue 22 SeptComputation and Language
The gist
Speech translation systems often struggle because they are trained on perfect transcripts but must work with imperfect speech transcriptions in real use. The authors propose a new way to train these systems by jointly improving how they recognize spoken words and translate them, using a technique that rewards good outcomes. Their tests with a multilingual model showed better translation quality and fewer recognition errors compared to traditional training methods. This approach helps reduce the mismatch between training conditions and real-world use, making speech translation more reliable.
Open → 2609.26536v1