Papers for

mental health technology builders

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multimodal dataset captures humour styles and emotions in video actors

MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions

Abstract: Computational recognition of verbal humour remains a challenging task, requiring an understanding of language, delivery style, emotions, and cultural context. Most existing approaches focus on binary classification and lack datasets that capture psychological dimensions of humour alongside variations in expression. We introduce MultiHuSE, a multimodal dataset comprising 2,407 high-definition videos of 50 demographically diverse actors performing 1,463 text samples across four psychological humour styles (affiliative, aggressive, self-enhancing, and self-deprecating), as well as neutral content. A subset is additionally annotated for underlying emotions. The dataset uniquely captures multiple actor interpretations of the same texts, enabling systematic analysis of expressive diversity. Baseline experiments show that multimodal fusion outperforms unimodal approaches (80.1% vs. 77.4% accuracy) in humour style classification, with particularly strong gains for affiliative humour (66% to 74%). While text provides the strongest individual signal, fusion models deliver meaningful improvements. We hope that MultiHuSE provides empirical support for psychological theories linking humour and emotion, while also opening new avenues for research in human communication, well-being, and AI-driven interaction. The dataset is available for academic use under an End-User Licence Agreement.

Thu 10 SeptComputation and LanguageComputer Vision and Pattern RecognitionMultimedia
The gist
Understanding humor by computers is hard because it involves language, emotions, and how something is said. The authors created a large video dataset called MultiHuSE with many actors showing different types of humor styles using the same texts. This dataset also includes emotions related to the humor and shows how the same joke can be expressed differently. They found using video, audio, and text together helps better identify humor styles than using just one type of information.
Open 2609.11322v1