EvolvingAvatar improves 3D head motion during conversations
EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold
Computer Vision and Pattern Recognition
Summary
Making 3D animated heads that talk and listen naturally is hard because they need to move their faces in sync with the conversation. The authors created EvolvingAvatar, a system that learns and adapts during a chat using the video and sound it sees and hears, even without pre-labeled examples of how to move. It adjusts its behavior over time within a conversation, improving how realistic the head motion looks and making it better match the person it is imitating. They also made a large set of conversation videos to test these systems and showed that EvolvingAvatar improves over time as the conversation goes on.
What this means in practice
- •For virtual reality developers: Create more natural and expressive 3D avatars that adapt their facial movements during real-time conversations.$Commercial implications: Enables VR platforms to offer dynamically adapting avatar animations, enhancing user immersion and communication quality.
- •For video game developers: Generate interactive character faces that adjust facial expressions and speech movements on-the-fly in response to player interactions.
Authors
Junjie Chen, Fei Wang, Kun Li, Yiqi Nie, Xun Yang, Yanbin Hao, Linfeng Zhang, Meng Wang
Abstract
Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAvatar, a causal generator that uses test-time training to adapt to user face video and dyadic audio during interaction. Its dyadic context prediction objective provides a self-supervised learning signal from audiovisual context without target motion labels at test time. Persistent fast weights accumulate these updates within each conversation to guide motion generation, while transient jaw adaptation responds to current audiovisual context. Predicted speech activity controls how persistent adaptation guides motion. We also introduce InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos. Experiments show improved conversational motion statistics over strong baselines. On the hardest out-of-distribution split, generation improves as conversations unfold, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.