SignRefine improves sign language video clarity with keypoint control

SignRefine: Adapting Foundational Video Models for Sign Language Generation

Computer Vision and Pattern Recognition

Summary

Making videos of sign language is tricky because the hand and facial movements must be very clear. Existing video models, which mostly learn from spoken-language videos, often make mistakes that make signs hard to understand. The authors created SignRefine, a model that uses simple 2D points showing hand and face positions to produce clearer sign language videos. They also introduced a new dataset called NVSign to help train their model on real sign language videos. Their model shows better accuracy in hand poses and is preferred by sign language users for how clear and understandable the videos are.

sign languagevideo generationvideo diffusion models2D keypoint conditioninghand pose accuracyfacial articulationtransformer modelsdatasetNVSignarticulation refinement

Authors

Anton Pelykh, Edward Fish, Ozge Mercanoglu Sincan, Richard Bowden

Abstract

Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions. Our approach builds on a pretrained video diffusion transformer and introduces local adapters with spatial grounding to selectively refine hand and face regions, steering the strong base model's prior toward accurate articulation. To enable this work and support broader sign language research, we present NVSign, a large-scale dataset of video content natively produced in sign language, offering diverse signer appearances, environments, and natural conversational settings. Trained on this data, our model shows up to 30% improvement in hand pose precision metrics over the strongest baseline and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.