Empathetic reinforcement learning adapts support to dialogue changes
Evolving Support Priorities in Empathetic Reinforcement Learning
Artificial Intelligence
Summary
People’s needs for emotional support change as conversations go on, but current AI methods use fixed rules that don’t adjust to these changes. The authors propose a new approach called CARE that changes how support is evaluated in real time based on the conversation’s context. CARE looks at three types of empathy—thinking, feeling, and taking action—to better help people during different stages of a dialogue. Their method improves AI performance in tests by adjusting what kind of support is rewarded depending on how the conversation unfolds and how the user feels.
What this means in practice
- •For customer support teams: Enhance chatbots to provide context-sensitive empathetic responses by adapting support priorities during customer interactions.
- •For mental health app developers: Develop virtual assistants that adjust empathetic support dynamically based on user emotions and dialogue progression for better engagement.
Authors
Pengyu Huang, Zhiyuan Han, Wenwen Tong, Hewei Guo, Jiangnan Chen, Sirui Chen, Lewei Lu, Beier Zhu, Xun Yang
Abstract
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.