Reactive audio-driven model improves listener facial motion in robots

REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

Robotics

Summary

Generating natural facial reactions for virtual or robotic listeners during conversations is tricky because listener responses depend on timing cues and subtle expressions like blinks. The authors created a system called REALM that first predicts broad listener facial movements and then adds small, random expression details triggered by the speaker's audio. Their model helps robots and avatars react more naturally to the speaker, making conversations feel more lifelike. They tested it on datasets and a humanoid robot to show it works better than previous methods.

What this means in practice

  • For robotics developers: Create humanoid robots with more natural and responsive facial expressions during conversations using an audio-driven reactive motion model.
  • For virtual assistant designers: Enhance avatar-based virtual assistants with realistic listening expressions to improve user engagement and communication clarity.

Authors

Peizhen Li, Longbing Cao, Yang Zhang

Abstract

Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ