Sprout builds and updates video memory dynamically during questioning

Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding

Computer Vision and Pattern Recognition

Summary

Long videos are hard for AI to understand all at once because they contain too much information. Existing methods first build a memory of the whole video and then try to answer questions, which is slow and wastes effort. The authors created Sprout, a system that watches videos in chunks and builds a memory while answering questions, saving time and updating what it knows as new questions come. Sprout stores video details as a tree of text summaries, keeping memory small and allowing the AI to revisit important parts more deeply when needed. This approach performs as well as or better than earlier methods but with less cost and no big setup step.

What this means in practice

  • For video analysis teams: Answer multiple questions about long videos efficiently by building memory dynamically during analysis, reducing resource use without losing accuracy.
  • For surveillance operators: Improve real-time video monitoring by updating event memories on the fly as inquiries come, allowing faster and more focused video review.

Authors

Wei Chen, Xuanyu Zheng, Yancheng Long, Haoyang Xu, Kaiyu Jiang, Bin Wen, Tingting Gao, Han Li, Long Chen

Abstract

Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and this pipeline is costly at both ends: with few questions, building memory for the whole video costs far more than answering them; with many questions, the memory is never updated, so what is learned while answering questions is lost to the next question. To alleviate these, we introduce Sprout, an agentic framework that builds memory while reasoning: a temporal tree that sprouts detailed nodes as questions are answered. The agent watches the video segment by segment at a low frame rate, stopping when the current question can be answered, remembers each segment as a coarse node of the tree, and revisits key intervals at a higher frame rate to refine the tree with the recovered details. Once a segment is recorded as text, its video input is removed from the context history, while the original video remains reachable through the video tools. The memory tree and prior question--answer records persist across questions, so the memory is online and dynamic: built from the first question onward and updated by every question thereafter. We find that replacing accumulated video inputs with textual memory substantially reduces context usage while maintaining accuracy, with slight improvements in some settings. Across benchmarks on three models, Sprout achieves competitive or improved accuracy relative to representative offline memory methods, with no upfront construction stage and lower context cost per question.