MetaSampling reduces frames for better long video question answering

MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering

Computer Vision and Pattern Recognition

Summary

Answering questions about very long videos can be hard because a computer has to look at many video frames to understand what's happening. The authors created MetaSampling, a method that helps choose fewer important frames to show the computer without losing accuracy in the answers. MetaSampling works on top of other frame selection methods and can sometimes even improve how well the computer answers questions. It was tested with many different video question answering systems and consistently reduced the number of frames needed.

What this means in practice

Authors

Ashim Dahal, Bikramjit Banerjee

Abstract

Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-$k$ embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We introduce \textbf{MetaSampling}, a training-free, plug-and-play sampling strategy that can be applied on top of existing frame selectors. MetaSampling improves downstream VQA efficiency by dynamically reducing the number of frames passed to the MLLM while preserving, and in some cases improving, answer accuracy. We evaluate MetaSampling across 36 paired frame-selector--MLLM-backbone--VQA-benchmark configurations. MetaSampling reduces the number of selected frames in all 36 configurations and improves accuracy in 25 of them, yielding an average frame reduction of $8.9\%$ while slightly improving accuracy overall.