Video sentiment analysis improves by splitting polarity and intensity tasks

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

MultimediaComputation and Language

Summary

Video sentiment analysis tries to understand feelings in videos using text, sound, and images together, but it is tricky because feelings have degrees and different parts work differently. The authors treat the problem as two tasks: recognizing positive or negative feelings and how strong they are. They introduce a method that uses small focused parts called bottlenecks to learn useful signals separately for these tasks from each kind of data. This helps the model better combine the different signals to predict feelings more accurately and understand subtle emotions in videos.

What this means in practice

  • For video content analysts: Improve emotion recognition in videos by separating sentiment polarity and intensity for better targeted analysis.
  • For multimedia app developers: Enhance apps that analyze user sentiment from videos by integrating compact, task-specific representations from text, audio, and visuals.

Authors

Ronghao Lin, Qiaolin He, Zefeng Lu, Yichu Liu, Li Huang, Sijie Mai, Haifeng Hu, Yap-peng Tan

Abstract

Video-based Multimodal sentiment analysis (MSA) must handle information from text, audio, and image sequence in human speaking videos, yet current methods often fail to integrate modalities with task awareness. Most models treat video sentiment prediction as a single task, overlooking its ordinal nature, and their fusion strategies struggle to capture diverse unique and synergic cues across modalities. To address these limitations, we adopt a divide-and-conquer perspective by reformulating MSA as an ordinal regression problem and decoupling it into polarity recognition and intensity prediction. Driven by information theory, we introduce a Mixture-of-Bottleneck (MoB) framework that assigns different latents to polarity- and intensity-specific experts for different modalities. With the learning of information bottleneck, each expert learns compact and task-relevant representations while filtering out redundancy and noise. A multimodal bottleneck routing fusion module then fuses these expert latents with hard mining strategy, guiding the prediction in the ordinal sentiment space. Extensive experiments on 4 MSA datasets and 4 language models show that MoB effectively leverages informative latents from diverse modalities and captures general sentiment structure. Beyond stronger performance, MoB comprehensively captures fine-grained intra- and inter-modal dynamics, enabling more trustworthy localization of nuanced video sentiment signals.