LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
2026-07-01 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionArtificial Intelligence
AI summaryⓘ
The authors created a new test called LongVQUBench to better evaluate how well large vision-language models understand video quality over long videos. Current tests mostly focus on short clips and simple issues, but this benchmark uses over 1200 videos with questions that check understanding at different levels, from spotting small problems to judging overall quality. They also added tricky tasks involving small, hard-to-find distortions to test detailed reasoning. Their experiments show that existing models struggle as videos get longer and require more complex thinking. The authors hope their work will help improve ways to measure long-term video quality understanding in these models.
LongVQUBenchVision-Language ModelsVideo Quality AssessmentTemporal ContinuityPerceptual ReasoningDistortion DetectionMultiple-Choice QuestionsOpen-Ended QuestionsLong-Term Video UnderstandingNeedle Distortion Question-Answering
Authors
Arpita Nema, Hanwei Zhu, Xi Zhang, Weisi Lin
Abstract
The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, and reasoning complexity inherent in long-duration content. To address these limitations, we present LongVQUBench, a comprehensive benchmark for long-term video quality understanding. LongVQUBench contains over 1200 diverse videos spanning movies, documentaries, surveillance footage, egocentric recordings, and animated content, accompanied by 1500 multiple-choice and open-ended questions for validation and testing. To assess perceptual reasoning across different temporal scopes, we introduce three progressively complex evaluation levels: (i) local event quality understanding (LQU) for analyzing localized distortions; (ii) cross-event quality reasoning (CQR) for integrating multiple degraded events; and (iii) global quality understanding (GQU) for holistic perceptual evaluation over extended durations. Furthermore, a needle distortion question-answering (NDQA) paradigm is embedded across all three levels, where spatial or temporal artifacts are sparsely inserted to probe fine-grained detection and reasoning capabilities. Extensive experiments on 14 state-of-the-art LVLMs reveal significant performance degradation with increasing video length and reasoning depth, highlighting their limited capacity for long-range temporal integration and perceptual attribution. We envision LongVQUBench as a foundational step toward the systematic, hierarchical, and explainable evaluation of LVLMs' long-term video quality understanding.