Multimodal emotion recognition improves with adaptive fusion and facial geometry

Multimodal Emotion Recognition in Conversations via Class-Wise Adaptive Modality Fusion and Affective Geometry

Computer Vision and Pattern Recognition

Summary

Understanding emotions in conversations is hard because people express feelings through words, tone, and facial expressions. The authors improve emotion recognition by combining facial appearance and precise facial movements, adjusting how much each type of signal counts depending on the emotion, and using a model of how emotions change over time. They tested their approach on two conversation datasets and found better accuracy, especially when emotions shift during the conversation. This shows that using detailed facial cues and emotion-aware mixing of signals helps machines understand feelings more reliably.

What this means in practice

  • For multimedia software developers: Create conversational agents that better detect user emotions by integrating facial movement patterns with audio and text cues.
  • For affective computing engineers: Build emotion recognition systems that dynamically adjust how visual, audio, and text signals are weighted depending on the expressed emotion.

Authors

Oriol Marín, Roger Marí, Gloria Haro, Rafael Redondo

Abstract

Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational context and emotional dynamics. We extend the Self-Distillation Transformer architecture for ERC with appearance+geometry visual representations, class-wise adaptive modality fusion, and a valence-arousal prior for affective transitions. On the MELD and IEMOCAP datasets, geometry-enhanced visual representations improve weighted F1 by 0.27 and 4.36 points over appearance-only features, respectively, while class-wise adaptive fusion provides further gains of 0.17 and 0.25 points over the original softmax gate. The valence-arousal prior yields targeted improvements of 0.30 and 0.74 accuracy points on emotionally shifted utterances while preserving performance on stable turns. These results indicate that structured facial cues, emotion-dependent modality weighting, and affective geometry provide complementary benefits for multimodal ERC.