SignFLIP enables smooth text and sign language conversion bidirectionally
SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale
Computation and LanguageComputer Vision and Pattern RecognitionMultimedia
Summary
Converting between sign language and spoken language is challenging because they use very different ways to express meaning. The authors created SignFLIP, a system that handles translating from sign language to text and generating sign language from text using the same model. It improves this process by training in stages on large amounts of data, refining a shared understanding of signs and words. SignFLIP matches or does better than specialized models for each task and can also help recognize sign language better.
What this means in practice
- •For speech recognition developers: Build apps that convert sign language videos into text and vice versa using a unified model with improved accuracy.
- •For mobile application developers: Create sign language communication tools that generate sign language videos from text messages and translate sign input back into text.$Commercial implications: This enables consumer apps for inclusive communication by integrating translation and generation capabilities into a single efficient system.
Authors
Zhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai, Rei Kawakami
Abstract
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.