Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots
2026-08-17 • Human-Computer Interaction
Human-Computer InteractionComputation and LanguageRobotics
AI summaryⓘ
The authors created a social robot called AffectLoop that listens and watches how a person’s emotions change during a conversation, then responds with both words and body language showing empathy. Unlike older systems that only read text and respond based on the user’s feelings at one moment, this robot considers ongoing emotional changes for both the speaker and itself. They tested the robot with five people and found it made users feel more understood and satisfied compared to a simpler version. The study shows that paying attention to emotional back-and-forth between a person and a robot can help the robot seem more empathetic.
empathetic dialogue systemsaffective dynamicsmultimodal interactionspeaker-listener modelverbal and facial affectLLM-based response generationembodied behaviorvalence-based distress recoveryMisty II robotaffective alignment
Authors
Zi Haur Pang, Casey Kennington, Tatsuya Kawahara
Abstract
Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primarily model empathy as a one-way mapping from the user's emotion to the system response, limiting their ability to capture embodied speaker--listener affective exchange. We present AffectLoop, a multimodal speaker-listener emotion-dynamics-aware spoken dialogue system implemented on the Misty II robot. The system tracks the speaker's verbal and facial affective dynamics, estimates the robot listener's own verbal and behavioral affective state, and conditions LLM-based response generation on both affective streams. The robot then generates a short spoken empathetic response together with emotionally congruent embodied behavior, forming a closed speaker--listener affective loop. We evaluate the system in a pilot within-subject study with five participants, comparing it with an otherwise identical utterance-conditioned baseline that omits the speaker- and listener-affective-state inputs. The proposed system received higher overall impression ratings, especially for empathetic response and user satisfaction. Post-hoc log analysis further showed higher speaker-listener affective alignment and stronger valence-based distress recovery. These preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.