Generate realistic listener reactions for natural video conversations

GLARE: Generating Listening Heads with Appropriate Reactions

Computer Vision and Pattern Recognition

Summary

People talking to each other often react with head nods, smiles, or surprised looks, but making computer-generated videos that show these real reactions is hard. The authors created a big new dataset of videos showing listeners’ reactions and labeled exactly when and how they react. They built a system called GLARE that listens to speech sounds and generates matching facial reactions, like nodding or laughing, at the right times. They also made new ways to check if these reactions happen naturally and look right. Their results show it's important to teach computers when and how to react, not just to make faces that look real.

What this means in practice

  • For virtual assistant developers: Create realistic listening behaviors in virtual agents that respond naturally during spoken interactions.$Commercial implications: Enables more engaging virtual assistants that can visually react in real time, improving user experience and marketability.
  • For video game animators: Add authentic listener reactions to non-player characters during in-game dialogues to boost realism and immersion.

Authors

Zikai Liao, Yumin Suh, Yi Ouyang, Yi-Lun Lee, Yi-Hsuan Tsai, Zhaozheng Yin

Abstract

While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.