Emotion concepts align across vision language models and data types
Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?
Computer Vision and Pattern RecognitionMachine Learning
Summary
The paper explores whether different vision-language AI models understand emotions in the same way across pictures, text, and model designs. The authors built a set of stories, real and fake images, and scenes to test this. They found that emotional meanings in images and text share a similar structure, and models can translate emotions between the two. Also, even different AI designs represent emotions with shared patterns that can be aligned to work together. This suggests that AI systems can have a common understanding of emotions, despite their differences.
What this means in practice
- •For ai product developers: Create AI tools that interpret and respond to emotions consistently across text and images using aligned model representations.$Commercial implications: Enables emotionally aware AI products integrating different data types with consistent responses.
- •For multimodal ai researchers: Design models and analyses that leverage shared emotional structures across modalities to improve generalization and interpretability.
Authors
Bohao Xing, Xin Liu, Kaishen Yuan, Deng Li, Rong Gao, Guoying Zhao, Xiaolan Fu, Heikki Kälviäinen
Abstract
Recent studies suggest that large language models encode emotion concepts as structured internal representations, but most existing work focuses on text and a single architecture. Therefore, we ask, do emotion concepts generalize across sources, modalities, and architectures in vision--language models (VLMs)? To address this, we construct CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes. For each stimulus source, we extract a separate set of six Ekman emotion vectors from each of three VLMs. We report four main findings as follows: 1) Image-derived emotion vectors form a low-dimensional geometry similar to that of text-derived vectors. Valence is relatively stable across sources, while arousal varies more. 2) Text- and image-derived emotion vectors have modest cosine similarity but still show held-out cross-modal correspondence. Text-derived vectors can also steer image interpretation. 3) Cross-architecture correspondence remains even when native cosine is near zero. Transformations estimated from generic ImageNet activations recover both correspondence and causal transfer without using the six emotion vectors or their labels. 4) After aligning representations across architectures, we construct a shared emotion subspace that preserves affective geometry and selective steering effects. The corresponding consensus emotion vectors also generalize to a held-out fourth architecture at two model sizes. These results suggest that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even when individual vector directions differ.