Prefix-Steered memory improves long video understanding with limited frames
PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Understanding very long videos is challenging because computers can only analyze a limited number of frames at once. The authors introduce a new memory method called PREM that summarizes visual information into a small, reusable memory. This memory helps answer questions about the video without repeatedly processing all frames or adding extra data. Their method improves accuracy on several video tasks using less computational resources than previous approaches.
What this means in practice
- •For video analysis engineers: Deploy efficient video systems that answer complex questions from long videos using limited frame inputs and small memory footprints.
- •For streaming video platform developers: Improve real-time understanding of live video streams with minimal additional processing and memory overhead.
Authors
Siru Zhong, Qiongyan Wang, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang
Abstract
Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches. We introduce Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs). PREM separates video ingestion from query answering: a recurrent writer distills visual streams into a compact 256 KiB multi-slot associative state, while a question-conditioned readout adds memory-derived key/value (K/V) steering modulations to existing non-visual prompt prefixes during prefill. This enables write-once, query-many inference without extra prompt tokens or decoding recurrence. Across six long-video benchmarks in offline and streaming end-of-stream settings, PREM consistently outperforms frozen baselines at every evaluated visual budget. Under a constrained budget of 16 frames, PREM improves macro-average accuracy by 3.06% on Qwen2.5-VL-3B, with gains of 11.0% on action antonym identification and 9.9% on localized needle retrieval. These gains require tuning 0.24% of backbone parameters at 0.03 GiB of peak GPU memory overhead.