Gesturefar enables real-time natural gesture generation from streaming speech

GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression

Computer Vision and Pattern RecognitionGraphicsHuman-Computer Interaction

Summary

Making digital characters move their hands naturally while talking is tricky, especially when the speech is still happening. Previous methods simplified complex hand motions into limited building blocks, which made the gestures less lifelike and varied. The authors created GestureFAR, a new way that predicts smooth, continuous hand movements directly from what is being said, without waiting for the speaker to finish. This approach lets virtual agents produce more natural gestures instantly as speech flows.

What this means in practice

  • For interactive game developers: Create game characters that produce natural hand gestures in real time as players speak or narrate.$Commercial implications: GestureFAR enables immersive, real-time character animation from live speech inputs, enhancing player experience and allowing commercialization of interactive storytelling.
  • For virtual assistant designers: Integrate real-time natural gestures with voice in conversational agents to improve user engagement and communication.

Authors

Pinxin Liu, Haiyang Liu, Jiahao Luo, Junhua Huang, Chunhao Zou, Luchuan Song

Abstract

Generating natural co-speech gestures from streaming speech is essential for embodied conversational agents, where motion must be produced while a user is still speaking. Recent streaming gesture systems make online generation possible by autoregressing over discrete motion tokens, but this design compresses high-dimensional continuous motion into finite codebooks and can limit the realism and diversity of generated gestures. To preserve both causality and continuous expressiveness, we propose \textbf{GestureFAR}, a flow-autoregressive framework for streaming co-speech gesture generation. First, GestureFAR autoregresses over causal continuous motion latents, using a transformer to model streaming audio-motion context and a per-token flow-matching head to sample the next latent from a continuous distribution. Second, we introduce a head-only flow distillation strategy that freezes the causal backbone and distills the multi-step per-token flow head into a single network evaluation using consistency and distribution-matching objectives. This keeps the model token-causal while removing the main latency bottleneck for live interaction. Experiments on BEAT2 show that GestureFAR significantly improves the quality--latency trade-off among streaming-capable methods, preserving strong gesture quality while enabling real-time token-causal generation. Project Page: https://andypinxinliu.github.io/GestureFAR