SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark

2026-07-20Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors created a new large dataset called SpEmoC to help computers recognize emotions in spoken conversations from movies and TV shows. Their dataset includes balanced examples of seven emotions with synchronized video, audio, and text, and uses careful splitting to avoid overlap between training and testing data. They show that having balanced data and strict dataset design improves how well emotion recognition models work, especially when applied to other datasets. Their work emphasizes the importance of good dataset construction for building more reliable emotion recognition systems.

Affective computingEmotion recognitionMultimodal dataDataset splittingClass imbalanceCross-dataset generalizationPretrained modelsHuman validationSpeech emotionBenchmarking
Authors
Sania Bano, Shahzad Ahmad, Santosh Kumar Vipparthi, Sukalpa Chanda, Subrahmanyam Murala
Abstract
Understanding human emotions in spoken conversations is a key challenge in affective computing, with applications in empathetic AI, human computer interaction, and mental health monitoring. However, existing datasets vary in scale, emotion distribution, modality alignment, and data partitioning strategies, which can influence reliable cross-dataset generalization and minority-emotion modeling. We introduce SpEmoC a Speaking segment Emotion for Conversations comprising 306,544 raw clips from 3,100 English language movies and TV series. From these, 30,000 high quality, class balanced clips are curated, featuring synchronized visual, audio, and textual modalities annotated for seven emotions through a hybrid pipeline that integrates pretrained models with human validation. SpEmoC uses strict movie- and series-level splits to prevent content overlap between split sets, allowing more reliable evaluation of model generalization. The dataset also maintains a near-balanced distribution across seven emotions, including minority classes such as Fear and Disgust, which supports more balanced learning across categories. Extensive experiments, including in-domain benchmarking, cross-dataset transfer, low-data training, class-imbalance analysis, and modality transfer show that balanced data and careful splitting lead to more stable performance across emotions when models are evaluated on other datasets. These results highlight the importance of dataset design for robust and transferable multimodal emotion recognition.