Psychology-based model improves spoken empathetic dialogue responses
ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models
Artificial Intelligence
Summary
Spoken dialogue systems often struggle to respond empathetically because they only mimic emotions instead of regulating them properly. The authors propose a new framework called ER-EDF that separates understanding a user's emotion from deciding how to express empathy in the response. This approach is based on psychological theories and works with existing large audio-language models. Their tests show that ER-EDF makes the system's responses more empathetic, both by automated scores and by humans judging them.
What this means in practice
- •For voice assistant developers: Improve voice assistants to respond more thoughtfully by regulating emotional responses rather than just imitating user feelings.
- •For customer service chatbot teams: Enhance automated spoken support services to deliver empathetic replies that adapt to users’ emotions more appropriately.
Authors
Hongyu Jin, Wenda Zhang, Runqiu Fei, Gongping Huang, Mike Conway, Ting Dang
Abstract
Empathetic response generation in spoken dialogue systems requires both accurate emotion perception and appropriate emotion regulation. Grounded in psychological theories such as the Perception-Action Model and emotion regulation theory, effective empathy depends not only on inferring a user's affective state but also on regulating how it is expressed in responses. However, recent large audio-language models (LALMs) largely treat emotion as a direct conditioning signal, lacking explicit regulatory mechanisms, which often leads to affect mirroring rather than calibrated support. We propose ER-EDF, a psychology-grounded framework that explicitly decouples emotion perception and emotion regulation in LALMs. Perception tracks the user's emotional state, while regulation determines how this state should guide empathetic response generation. The framework is model-agnostic and integrates seamlessly into existing LALMs. We further construct a spoken empathetic dialogue dataset and introduce empathy-aware evaluation metrics beyond lexical matching. Experiments across five LALMs and two datasets show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.