BranchShine-CR improves multilingual phonetic transcription accuracy efficiently

BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization

Machine Learning

Summary

Transcribing many languages into a universal set of speech sounds called IPA is hard, especially on small devices. The authors introduce BranchShine-CR, a compact machine-learning model that understands many languages and transcribes their sounds more accurately than previous models with fewer parameters. It uses a combination of new techniques to stay accurate while being lightweight. This makes it easier to add detailed speech recognition into small devices even when computing resources are limited.

What this means in practice

  • For mobile app developers: Integrate efficient IPA transcription for speech-based language learning and pronunciation feedback on devices with limited computing power.
  • For assistive technology engineers: Build compact speech recognition tools for multilingual and low-resource contexts to improve accessibility features in wearable and embedded devices.

Authors

Nikhil Navas, Sergio Chevtchenko, Talisson Damiao, Saeed Afshar

Abstract

We introduce BranchShine-CR, a 25M-parameter model for multilingual transcription into the International Phonetic Alphabet (IPA). It combines log-mel features, a rotary-position E-Branchformer encoder, intermediate self-conditioned connectionist temporal classification (CTC), and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances, it achieves 4.47% IPA character error rate, a 22.3% relative reduction from ZIPA-CTC-NS, with approximately one-twelfth as many parameters while being trained from scratch. BranchShine-CR also outperforms a similarly sized NeMo Conformer baseline across all 41 dataset language labels. Ablation studies indicate the individual components synergetically acting in model performance contribution. These findings support compact IPA recognition capabilities under limited compute budget, for applications in low-resource on-device pronunciation assessment.