Simultaneous translation enables live conversation across sign languages
Simultaneous Translation between Sign Languages
Computer Vision and Pattern RecognitionComputation and Language
Summary
People who are deaf or hard of hearing often use different sign languages, making real-time conversations difficult. Existing translation systems need to see a whole sentence before translating, causing delays. The authors created a new system that translates sign languages while the person is still signing, reducing delays and supporting live conversations. They also developed a way to measure how quickly and accurately the system works. Their method balances speed and accuracy, even when sign languages have different word orders.
What this means in practice
- •For video call platforms: Offer real-time sign language translation during two-way video calls to enable smoother communication between users of different sign languages.$Commercial implications: Provides a live sign-to-sign translation service for video conferencing companies to improve accessibility for deaf users worldwide.
- •For broadcast interpreters: Integrate streaming sign language translation to provide live sign language interpretation across different sign languages during televised events.
Authors
Zetian Wu, Bowen Xie, Stefan Lee, Liang Huang
Abstract
Deaf and hard-of-hearing (DHH) signers cannot converse in real time across different sign languages today: existing sign-to-sign translation systems run offline, requiring the full source clip before any target sign is emitted. Live use cases - e.g. broadcast interpretation and two-way video calls - instead demand simultaneous output, while the source signer is still signing. We present, to our knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision. We further introduce ca-Stream-AL, a computation-aware latency metric for streaming output. Averaged across six S2S directions on both a smaller human-verified test set and a larger synthetic S2S corpus, our streaming system achieves a 38% ca-Stream-AL reduction while staying within a 9% DTW-PA-MPJPE increase and a 2.1 BLEU-4 drop compared to the full-sentence baseline. A word-order case study probes how the streaming model handles word order mismatch between different sign languages - a consequence of simultaneous translation.