Large AI models struggle to understand cinematic storytelling techniques
CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models
Computer Vision and Pattern Recognition
Summary
Understanding how movies tell stories visually is hard for AI models that see and read videos. The authors created a new test called CinematicVQA, which checks if AI can reason about the film-making methods behind what we see, not just describe the images. They found that AI models are better at describing scenes visually than explaining the storytelling tricks used, and common techniques to improve reasoning did not help. Training AI on this new test helped them get better at understanding narrative and multi-step story reasoning.
What this means in practice
- •For video editing teams: Identify and improve storytelling impact in video projects using AI trained to reason about cinematic techniques.
- •For media content platforms: Enhance video recommendation systems by integrating AI that understands film grammar and narrative structure.
Authors
Shuo Xing, Pooja Verlani, Balu Adsumilli, Zhengzhong Tu
Abstract
Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.