Papers for

multimedia app developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Video sentiment analysis improves by splitting polarity and intensity tasks

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

Abstract: Video-based Multimodal sentiment analysis (MSA) must handle information from text, audio, and image sequence in human speaking videos, yet current methods often fail to integrate modalities with task awareness. Most models treat video sentiment prediction as a single task, overlooking its ordinal nature, and their fusion strategies struggle to capture diverse unique and synergic cues across modalities. To address these limitations, we adopt a divide-and-conquer perspective by reformulating MSA as an ordinal regression problem and decoupling it into polarity recognition and intensity prediction. Driven by information theory, we introduce a Mixture-of-Bottleneck (MoB) framework that assigns different latents to polarity- and intensity-specific experts for different modalities. With the learning of information bottleneck, each expert learns compact and task-relevant representations while filtering out redundancy and noise. A multimodal bottleneck routing fusion module then fuses these expert latents with hard mining strategy, guiding the prediction in the ordinal sentiment space. Extensive experiments on 4 MSA datasets and 4 language models show that MoB effectively leverages informative latents from diverse modalities and captures general sentiment structure. Beyond stronger performance, MoB comprehensively captures fine-grained intra- and inter-modal dynamics, enabling more trustworthy localization of nuanced video sentiment signals.

Wed 16 SeptMultimediaComputation and Language
The gist
Video sentiment analysis tries to understand feelings in videos using text, sound, and images together, but it is tricky because feelings have degrees and different parts work differently. The authors treat the problem as two tasks: recognizing positive or negative feelings and how strong they are. They introduce a method that uses small focused parts called bottlenecks to learn useful signals separately for these tasks from each kind of data. This helps the model better combine the different signals to predict feelings more accurately and understand subtle emotions in videos.
Open → 2609.18470v1