Joint training improves speech recognition and translation accuracy
Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation
Computation and Language
Summary
Speech translation systems often struggle because they are trained on perfect transcripts but must work with imperfect speech transcriptions in real use. The authors propose a new way to train these systems by jointly improving how they recognize spoken words and translate them, using a technique that rewards good outcomes. Their tests with a multilingual model showed better translation quality and fewer recognition errors compared to traditional training methods. This approach helps reduce the mismatch between training conditions and real-world use, making speech translation more reliable.
What this means in practice
- •For speech technology developers: Implement joint reward training to improve transcription and translation accuracy in speech-to-text translation apps.
- •For multilingual customer support teams: Use enhanced speech translation systems for more accurate understanding and response across different languages in real-time calls.
Authors
Yanghe Dong, Wanting Huang, Weiran Wang
Abstract
In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimization (GRPO). We score both transcripts and translations, with translation conditioned on model-generated transcripts, and compare three token advantage strategies. Using Qwen2.5-Omni-3B across four languages, we evaluate CoT against direct speech translation (Direct ST) under SFT and GRPO, training on CoVoST 2 and testing on CoVoST 2 and FLEURS. CoT GRPO outperforms Direct ST GRPO by 1.77 and 0.83 average BLEU points on CoVoST 2 and FLEURS. Compared to CoT SFT, GRPO boosts BLEU by 0.82 and 0.67 points and reduces word error rate (WER) by 8.8% and 7.2% relatively. These results highlight reinforcement fine-tuning as an effective method to mitigate the training-inference mismatch, jointly improving recognition and translation.