Large language models improve real time speech translation with less delay

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

Computation and Language

Summary

Translating spoken language in real time is hard because it requires quickly understanding and converting speech while keeping the speaker’s voice. The authors developed a new way to train language models that respect the timing of speech and keep the speaker’s voice identity. This method uses less training data but still delivers better translation quality and faster responses in multiple languages. Their system balances accuracy and speed better than previous methods, helping conversations flow more naturally across languages.

What this means in practice

  • For real-time communication service developers: Create faster and more accurate live speech translation apps that keep the speaker's voice consistent across languages.$Commercial implications: Enables live multi-language communication platforms with improved quality and reduced delay, appealing to global conferencing and call services.
  • For multilingual customer support teams: Integrate speech translation that reduces waiting times and improves clarity in conversations with international clients.

Authors

Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu, Harishchandra Dubey, Myungjong Kim, Zhehuai Chen, Oluwatobi Olabiyi, Sanjeev Khudanpur

Abstract

Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.