Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

2026-08-24Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and LanguageInformation RetrievalMultimedia
AI summary

The authors created a new test called the Cultural Moment Benchmark (CMB) to better understand how well computers grasp cultural meanings in videos. They break down the task into three parts: naming the cultural concept, spotting it in a video, and pinpointing when it happens. Testing several AI models showed that even the best ones struggle, especially when needing all three parts correct, and that audio can sometimes help or confuse the recognition depending on the culture. Their human study also found that people unfamiliar with certain cultures find these tasks very hard, showing the challenge is specific to cultural knowledge. Overall, the authors use CMB to identify which part of understanding culture in video causes problems for machines.

Cultural understandingVideo recognitionTemporal localizationVision-language modelsCultural benchmarksSemantic similaritySub-event detectionCross-cultural knowledgeAudio-visual integrationSoutheast Asian culture
Authors
Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo
Abstract
Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this into three abilities: naming what a concept symbolizes, visually recognizing it on video, and locating its sub-events in time. Existing video-cultural benchmarks tend to test what is seen, collapsing these three abilities into a single score that hides the bottleneck. We introduce the Cultural Moment Benchmark (CMB): 306 expert-curated concepts from seven countries in Southeast Asia across five categories. We evaluate each concept through three stages, one per ability. Given a description, Stage 1 (S1) selects from four candidate concept names, Stage 2 (S2) selects from four candidate video moments, and Stage 3 (S3) predicts the start and end times of the moment in a video. To keep each stage focused on a distinct ability, we use three design choices: semantic-similarity distractors (S1, S2), unlabeled video moments (S2), and free-form localization on a different example video (S3). Across six vision-language models, failure modes vary by ability and modality. i) Even the strongest closed-source models score below 30% when all three stages must be correct; ii) The three abilities do not fully cascade: naming a concept correctly helps half the models recognize it on video, but recognizing it has little effect on locating the sub-event in time; iii) Audio is complementary, redundant, or distracting depending on the concept, more often distracting in non-Latin-script countries; removing both audio and subtitles hurts Games and Music the most. Our 14-rater human study shows that even Expert raters score below chance on concepts from a neighboring country, indicating that CMB requires country-specific cultural knowledge. CMB acts as a diagnostic harness, attributing failures to a specific ability or modality.