Dual-stream attention boosts real-time chinese to english translation quality
Dual-Stream Simultaneous Translation via 2D Grid Attention
Artificial Intelligence
Summary
Simultaneous translation means starting to translate before hearing the whole sentence. The authors designed a new method that looks at both source and target sentences together in a grid, allowing the system to better understand how pieces from each relate as translation happens. This method makes translation more accurate and faster by learning when to read or write without needing extra guidance. They tested it on Chinese to English translation and found it did better than previous methods while keeping the delay very low.
What this means in practice
- •For real-time translation system developers: Improve streaming translation systems by integrating dual-stream attention that balances accuracy and latency more effectively during live conversations.
- •For speech-to-text application teams: Enhance incremental text generation with adaptive read/write decisions using joint attention to reduce delay and improve output quality.
Tested on one dataset.
Authors
Yu Pu, Wei-Qiang Zhang
Abstract
Simultaneous machine translation must generate target tokens before the source input is complete. Existing approaches address this through post-hoc read-write policies, leaving the attention mechanism unaware of bidirectional stream dependencies. We propose a dual-stream attention framework that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization. Two approximations---broadcast and Hadamard---reduce the per-layer complexity from O(X^2Y+XY^2) to O(X^2+Y^2+XY) with provably decaying error. Training uses a self-guided loop: a per-cell loss heatmap drives dynamic-programming path recovery, which generates read/write decision supervision labels without external alignment. An incremental KV cache with anchored rotary position embeddings enables efficient streaming inference. On Chinese-to-English simultaneous translation, the proposed model outperforms the Wait-k baseline by +5.66 BLEURT and +10.36 COMET at comparable latency, and surpasses the non-streaming reference on COMET at a fraction of the response delay.