Joint training improves speech recognition and translation accuracy

Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation

Computation and Language

Summary

Speech translation systems often struggle because they are trained on perfect transcripts but must work with imperfect speech transcriptions in real use. The authors propose a new way to train these systems by jointly improving how they recognize spoken words and translate them, using a technique that rewards good outcomes. Their tests with a multilingual model showed better translation quality and fewer recognition errors compared to traditional training methods. This approach helps reduce the mismatch between training conditions and real-world use, making speech translation more reliable.

What this means in practice

Authors

Yanghe Dong, Wanting Huang, Weiran Wang

Abstract

In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimization (GRPO). We score both transcripts and translations, with translation conditioned on model-generated transcripts, and compare three token advantage strategies. Using Qwen2.5-Omni-3B across four languages, we evaluate CoT against direct speech translation (Direct ST) under SFT and GRPO, training on CoVoST 2 and testing on CoVoST 2 and FLEURS. CoT GRPO outperforms Direct ST GRPO by 1.77 and 0.83 average BLEU points on CoVoST 2 and FLEURS. Compared to CoT SFT, GRPO boosts BLEU by 0.82 and 0.67 points and reduces word error rate (WER) by 8.8% and 7.2% relatively. These results highlight reinforcement fine-tuning as an effective method to mitigate the training-inference mismatch, jointly improving recognition and translation.