KeyRec reduces memory use for understanding long videos and streams

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Videos and streams are long, and computers often slow down trying to watch all the details at once. The authors designed KeyRec, a system that keeps important recent moments in short-term memory and groups earlier events in a smart way to save space. It decides how much attention to give to recent versus older parts based on the question asked, without rewatching old video parts. This makes video understanding faster and more efficient without losing important details.

What this means in practice

  • For video processing engineers: Develop real-time video analysis systems that handle long streams efficiently with less memory by selectively storing key visual information.
  • For surveillance system designers: Build surveillance tools that maintain crucial event memories and recent frames to answer queries about incidents without storing entire footage.

Authors

Zihan Chen, Xuejian Rong, Xiaojuan Wang, Boqing Gong, Adi Zicher, Yael Pritch, Nikhil Karnad

Abstract

Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose \textbf{KeyRec}, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.