Bangla memes challenge AI to understand culture and humor better

BanglaMemeX: Advancing Cultural Metaphoric Image Interpretation in Bangla with a Multimodal Explainable Dataset

Computation and LanguageComputer Vision and Pattern Recognition

Summary

AI models that combine pictures and words work well on many tasks but have trouble understanding jokes and meanings hidden in cultural references, like those in Bangla memes. The authors made a special collection of 3,000 Bangla memes labeled with emotions and explanations of the jokes and pictures. They tested current AI tools and found that while these tools can recognize obvious details, they struggle to grasp the deeper cultural meanings. This shows that AI needs to be improved to understand language and culture together, especially for less common languages like Bangla.

Vision Language Modelsmultimodal benchmarksBangla languageinternet memescultural symbolismsarcasm detectionmetaphor interpretationcode-mixingexplanation generation

Authors

Md. Sadman Sakib, Zisan Mahmud, Md. Fahim Arefin, Md Fahim

Abstract

Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason about culturally grounded and metaphor-rich content remains insufficiently studied. Internet memes present a challenging setting where meaning emerges from implicit interactions between image, overlaid text, sarcasm, and shared socio-cultural knowledge rather than literal visual recognition. This challenge is amplified in low-resource languages such as Bangla, where code-mixing, stylized scripts, and culturally specific symbolism introduce substantial distribution shift. In this work, we introduce BanglaMemeX, a culturally grounded multimodal benchmark comprising 3,000 Bangla memes annotated with multi-dimensional labels (humor, sarcasm, offensiveness, motivational intent, and overall sentiment) and human-written explanations that explicitly describe textual and visual metaphors. We systematically evaluate modern VLMs on both classification and explanation generation, revealing that current models struggle to interpret implicit cultural cues despite reasonable surface-level accuracy. Our results highlight the need for culturally-aware multimodal systems capable of grounded reasoning under linguistic and cultural distribution shift.