BranchShine-CR improves multilingual phonetic transcription accuracy efficiently
BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization
Machine Learning
Summary
Transcribing many languages into a universal set of speech sounds called IPA is hard, especially on small devices. The authors introduce BranchShine-CR, a compact machine-learning model that understands many languages and transcribes their sounds more accurately than previous models with fewer parameters. It uses a combination of new techniques to stay accurate while being lightweight. This makes it easier to add detailed speech recognition into small devices even when computing resources are limited.
What this means in practice
- •For mobile app developers: Integrate efficient IPA transcription for speech-based language learning and pronunciation feedback on devices with limited computing power.
- •For assistive technology engineers: Build compact speech recognition tools for multilingual and low-resource contexts to improve accessibility features in wearable and embedded devices.
Authors
Nikhil Navas, Sergio Chevtchenko, Talisson Damiao, Saeed Afshar
Abstract
We introduce BranchShine-CR, a 25M-parameter model for multilingual transcription into the International Phonetic Alphabet (IPA). It combines log-mel features, a rotary-position E-Branchformer encoder, intermediate self-conditioned connectionist temporal classification (CTC), and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances, it achieves 4.47% IPA character error rate, a 22.3% relative reduction from ZIPA-CTC-NS, with approximately one-twelfth as many parameters while being trained from scratch. BranchShine-CR also outperforms a similarly sized NeMo Conformer baseline across all 41 dataset language labels. Ablation studies indicate the individual components synergetically acting in model performance contribution. These findings support compact IPA recognition capabilities under limited compute budget, for applications in low-resource on-device pronunciation assessment.