Streaming video agent selects key frames without future knowledge
SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection
Computer Vision and Pattern Recognition
Summary
Selecting important frames from videos usually requires seeing the whole video first. The authors studied how to pick key frames from a video as it streams, without knowing how long it is or what questions might be asked about it. They built a system called SVMemAgent that keeps a small memory of important past frames and decides whether to update it with new frames on the fly. Their system learns from many question-answer examples so it can keep generally useful frames even when no question is known. This helps in tasks like video question answering by focusing on frames with important information.
What this means in practice
- •For video analytics teams: Integrate a frame selection method that works live on streaming video without needing full video or query access.
- •For surveillance operators: Improve monitoring systems by dynamically storing key frames from live video feeds to better support later querying.
Authors
Dohwan Ko, Ji Soo Lee, Pierce Chuang, Debojeet Chatterjee, Ashish Shenoy, Yichao Lu, Seungwhan Moon, Xin Luna Dong, Vikas Bhardwaj, Hyunwoo J. Kim
Abstract
Most keyframe selection studies focus on offline settings, assuming access to the full video and query in advance. In contrast, real-world streaming scenarios require online frame selection under unknown video duration, without access to either the query or future frames during selection. To address this, we introduce Streaming Video Memory (SVMem), a compact and representative memory of previously observed content, updated continuously as the video stream unfolds. Building on this setting, we propose the Streaming Video Memory Agent (SVMemAgent), which dynamically maintains a memory by deciding at each timestep whether to replace an existing memory frame with the incoming frame or discard it. SVMemAgent is trained using Group Relative Policy Optimization (GRPO) with task-driven rewards derived from diverse question-answer pairs, implicitly exposing the policy to a distribution of queries during training so that SVMem retains generally informative frames at inference, when queries are unavailable. Experiments on both online and offline video benchmarks show that SVMemAgent consistently outperforms online frame selection baselines and achieves competitive performance with offline methods that assume access to the full video and query. Through task-driven rewards, SVMemAgent learns an emergent keyframe selection policy that prefers frames containing textual information, which may benefit downstream VideoQA tasks.