CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

2026-08-03Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionComputation and LanguageMultimedia
AI summary

The authors created CultureVidBench, a test set to check how well text-to-video models can show different cultures in their videos. They included 1,000 prompts about various countries and cultural elements like traditions and visible text. They tested seven models and found that while the videos usually match the text well and look good, the models often miss important cultural details, especially for less-known traditions and sounds. This shows current video models need improvement in accurately representing diverse cultures.

text-to-video generationbenchmarkcultural representationmultimodalsemantic adherenceritualsvisible textaudio cuessocial interactionsuser studies
Authors
Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu
Abstract
Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice & performance, and ritual & ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM-based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.