Psychology-based model improves spoken empathetic dialogue responses

ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models

Artificial Intelligence

Summary

Spoken dialogue systems often struggle to respond empathetically because they only mimic emotions instead of regulating them properly. The authors propose a new framework called ER-EDF that separates understanding a user's emotion from deciding how to express empathy in the response. This approach is based on psychological theories and works with existing large audio-language models. Their tests show that ER-EDF makes the system's responses more empathetic, both by automated scores and by humans judging them.

What this means in practice

Authors

Hongyu Jin, Wenda Zhang, Runqiu Fei, Gongping Huang, Mike Conway, Ting Dang

Abstract

Empathetic response generation in spoken dialogue systems requires both accurate emotion perception and appropriate emotion regulation. Grounded in psychological theories such as the Perception-Action Model and emotion regulation theory, effective empathy depends not only on inferring a user's affective state but also on regulating how it is expressed in responses. However, recent large audio-language models (LALMs) largely treat emotion as a direct conditioning signal, lacking explicit regulatory mechanisms, which often leads to affect mirroring rather than calibrated support. We propose ER-EDF, a psychology-grounded framework that explicitly decouples emotion perception and emotion regulation in LALMs. Perception tracks the user's emotional state, while regulation determines how this state should guide empathetic response generation. The framework is model-agnostic and integrates seamlessly into existing LALMs. We further construct a spoken empathetic dialogue dataset and introduce empathy-aware evaluation metrics beyond lexical matching. Experiments across five LALMs and two datasets show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.