Papers for

real-time communication service developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Large language models improve real time speech translation with less delay

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

Abstract: Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.

Thu 24 SeptComputation and Language
The gist
Translating spoken language in real time is hard because it requires quickly understanding and converting speech while keeping the speaker’s voice. The authors developed a new way to train language models that respect the timing of speech and keep the speaker’s voice identity. This method uses less training data but still delivers better translation quality and faster responses in multiple languages. Their system balances accuracy and speed better than previous methods, helping conversations flow more naturally across languages.
Open → 2609.30416v1