Empathy in large language models is steerable but complex and multidimensional
Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs
Computation and Language
Summary
The paper studies how to control empathy shown by large language models (LLMs), breaking it down into different types like emotional reaction and interpretation. The authors found that while it is possible to adjust empathy, it isn’t controlled by just one simple switch but involves multiple interacting parts inside the model. They also discovered that changing the model's personality style affects empathy in ways that can’t be fully explained by targeting single directions in the model’s internal workings. This means controlling empathy in LLMs needs more nuanced techniques than simple adjustments.
What this means in practice
- •For chatbot developers: Adjust chatbot empathy responses more precisely by targeting multiple activation directions simultaneously within the model.
- •For customer service platform teams: Customize the personality-driven empathy of AI agents by combining persona prompts with targeted internal activation control for improved user interaction.
Authors
JuHeon Ha, Byounghan Lee, Yunseo Choi, Kyung-Ah Sohn
Abstract
Activation steering has been used to control traits such as honesty, refusal, and sycophancy, yet supportive empathy is evaluated along multiple dimensions that need not correspond to independently controllable activation directions. Using the EPITOME framework, which decomposes supportive empathy into Emotional Reactions, Interpretations, and Explorations, we study three instruction-tuned LLMs and ask whether candidate directions derived from these labels produce distinguishable intervention effects or instead share structure, and how persona prompts interact with those directions. We find that contrastive activation addition yields a stable middle-layer intervention that consistently shifts the EPITOME proxy scores across models, moving empathy analysis beyond response-level scoring. However, the recovered directions are only partially separable: steering one direction induces off-target shifts, and hand-crafted prompting shifts the empathy profile rather than isolating a single dimension. Persona prompts substantially change EPITOME scores, but a paired activation-shift decomposition shows that the recovered subspace captures only approximately 3 percent of persona-induced squared activation-shift magnitude at layer 15. Under this EPITOME-based definition, expressed empathy is steerable but multi-axial, and controlling persona-conditioned empathy requires targeting structure beyond individual mechanism directions.