Papers for

surveillance operators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Sprout builds and updates video memory dynamically during questioning

Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding

Abstract: Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and this pipeline is costly at both ends: with few questions, building memory for the whole video costs far more than answering them; with many questions, the memory is never updated, so what is learned while answering questions is lost to the next question. To alleviate these, we introduce Sprout, an agentic framework that builds memory while reasoning: a temporal tree that sprouts detailed nodes as questions are answered. The agent watches the video segment by segment at a low frame rate, stopping when the current question can be answered, remembers each segment as a coarse node of the tree, and revisits key intervals at a higher frame rate to refine the tree with the recovered details. Once a segment is recorded as text, its video input is removed from the context history, while the original video remains reachable through the video tools. The memory tree and prior question--answer records persist across questions, so the memory is online and dynamic: built from the first question onward and updated by every question thereafter. We find that replacing accumulated video inputs with textual memory substantially reduces context usage while maintaining accuracy, with slight improvements in some settings. Across benchmarks on three models, Sprout achieves competitive or improved accuracy relative to representative offline memory methods, with no upfront construction stage and lower context cost per question.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Long videos are hard for AI to understand all at once because they contain too much information. Existing methods first build a memory of the whole video and then try to answer questions, which is slow and wastes effort. The authors created Sprout, a system that watches videos in chunks and builds a memory while answering questions, saving time and updating what it knows as new questions come. Sprout stores video details as a tree of text summaries, keeping memory small and allowing the AI to revisit important parts more deeply when needed. This approach performs as well as or better than earlier methods but with less cost and no big setup step.
Open → 2609.35497v1

Streaming video agent selects key frames without future knowledge

SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection

Abstract: Most keyframe selection studies focus on offline settings, assuming access to the full video and query in advance. In contrast, real-world streaming scenarios require online frame selection under unknown video duration, without access to either the query or future frames during selection. To address this, we introduce Streaming Video Memory (SVMem), a compact and representative memory of previously observed content, updated continuously as the video stream unfolds. Building on this setting, we propose the Streaming Video Memory Agent (SVMemAgent), which dynamically maintains a memory by deciding at each timestep whether to replace an existing memory frame with the incoming frame or discard it. SVMemAgent is trained using Group Relative Policy Optimization (GRPO) with task-driven rewards derived from diverse question-answer pairs, implicitly exposing the policy to a distribution of queries during training so that SVMem retains generally informative frames at inference, when queries are unavailable. Experiments on both online and offline video benchmarks show that SVMemAgent consistently outperforms online frame selection baselines and achieves competitive performance with offline methods that assume access to the full video and query. Through task-driven rewards, SVMemAgent learns an emergent keyframe selection policy that prefers frames containing textual information, which may benefit downstream VideoQA tasks.

Wed 16 SeptComputer Vision and Pattern Recognition
The gist
Selecting important frames from videos usually requires seeing the whole video first. The authors studied how to pick key frames from a video as it streams, without knowing how long it is or what questions might be asked about it. They built a system called SVMemAgent that keeps a small memory of important past frames and decides whether to update it with new frames on the fly. Their system learns from many question-answer examples so it can keep generally useful frames even when no question is known. This helps in tasks like video question answering by focusing on frames with important information.
Open → 2609.18540v1