Video language models struggle to judge day long videos accurately
PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language
Summary
The paper looks at how well video-language models can judge and understand very long videos, like those lasting a whole day or more. The authors created a new testing system called PlaylistEval that automatically generates challenging questions across 100-hour video playlists without needing humans to label data. They found that even the best current models only correctly judge about 75% of the questions, and their accuracy drops as the video playlists get longer. Using multiple types of information, like audio and visuals, helps the models perform better.
What this means in practice
- •For video platform engineers: Test and improve automated video understanding systems on day-long content with PlaylistEval to better measure model accuracy without manual labeling.
- •For multimodal ai developers: Use PlaylistEval's evaluation framework to benchmark and enhance video-language judge models for tasks involving ultra-long video data.
Authors
Shayekh Bin Islam, Hwanjun Song
Abstract
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at https://playlisteval.github.io.