Vimarsha improves speech recognition testing for diverse Indian languages

Vimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations

Computation and Language

Summary

Automatic speech recognition (ASR) systems often give misleading results when tested on neat, controlled recordings or when judged by strict spelling rules that don't allow real language differences. The researchers created Vimarsha, a large and diverse set of speech recordings covering all 22 official Indian languages, including hard, real-world audio and multiple correct transcriptions for each sentence. Testing popular ASR models on Vimarsha showed big changes in which models performed best and revealed problems linked to speakers’ regions, backgrounds, speaking speed, and noise. This helps better understand how well speech technology works for real Indian language speakers.

What this means in practice

  • For speech technology developers: Test ASR models on diverse, real-world Indian language data including multiple valid transcriptions to better assess performance across demographics and conditions.
  • For call center technology teams: Improve language recognition accuracy by evaluating systems with challenging, varied Indian language audio resembling real customer calls.

Authors

Kaushal Santosh Bhogale, Srija Anand, Sadakopa Ramakrishnan Thothathiri, Tahir Javed, Sshubam Verma, Mitesh M. Khapra

Abstract

Evaluation benchmarks for Indian language automatic speech recognition (ASR) suffer from two systematic biases: optimistic scores from clean, controlled audio conditions, and pessimistic scores from overly rigid transcription standards that penalize valid linguistic variations. We introduce Vimarsha, a 100-hour benchmark spanning all 22 scheduled Indian languages, designed to address both distortions. Vimarsha combines demographically diverse on-field recordings with carefully mined in-the-wild audio selected for acoustic difficulty, alongside a lattice of variations framework that encodes multiple valid transcriptions per utterance. Evaluations of 10 state-of-the-art ASR models reveal substantial shifts in model rankings under realistic conditions, geographic and demographic performance disparities, and systematic failure modes across speaking rates and acoustic environments.