Query-guided summarization improves locating events in long videos
Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Finding the right moment in a long video that matches a spoken or written question is hard because videos have lots of extra, unimportant parts. The researchers created a new method that first summarizes parts of the video based on the question, then picks the most relevant summaries to find the exact moment. This reduces memory needs and helps the system focus better. Their method, called SumGround, works better than previous techniques, especially for long videos.
What this means in practice
- •For video content analysts: Automatically identify precise relevant moments in long video footage based on natural language queries with less memory use.
- •For security monitoring teams: Pinpoint specific events in continuous surveillance videos by processing condensed summaries guided by detailed incident descriptions.
Authors
Nanxing Hu, Xiaoyue Duan, Qiwei Yan, Kailin Lyu, Jinchao Zhang, Guoliang Kang
Abstract
Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query-relevant evidence. Instead of dense frame sampling which incurs prohibitive training memory, previous reinforcement learning with verifiable rewards (RLVR) works typically utilize sparse sampling, which makes training feasible but may miss critical evidence. In this paper, we propose a ``summarize before grounding'' framework (named ``SumGround'') for long-video temporal grounding. The key of SumGround is to perform query-guided chunk condensation to aggregate and retrieve query-relevant evidence. Specifically, we split the video into several chunks and perform two-level chunk condensation. First, we introduce query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries. Furthermore, we design an associative summary retrieval scheme to rank and select chunk summaries that are most likely to contain the event interval. Both query-guided latent summary and associative summary retrieval schemes are enabled by RLVR. To reduce memory consumption, we propose a length-aware gradient gating module to selectively stop gradient back-propagated to visual tokens. Extensive experiments demonstrate that SumGround performs favorably against previous state-of-the-art methods across multiple downstream datasets, with remarkable gains on long videos.