Space improves video memory by predicting future use patterns
SPACE: Sparse Predictive Attractor via Counterfactual Eviction for Streaming Video Memory
Computer Vision and Pattern Recognition
Summary
When computers watch videos and remember what they see over time, they have limited memory and need to decide what to keep or forget. The authors propose a way to predict how different choices about what to remember affect future video understanding. Their method, called SPACE, uses a special model to foresee future video parts that will be useful and makes smarter decisions about what to keep in memory. This helps keep important information longer and improves overall video performance without needing to relearn during use.
What this means in practice
- •For video processing engineers: Design memory systems in video pipelines to better retain information that will be useful for future video tasks under tight memory limits.
- •For autonomous vehicle developers: Improve how onboard systems remember and predict important future video scenes to enhance decision making in real-time driving environments.
Authors
Hongjin Niu, Weizhan Zhang, Shuo Bao, Jiahao Wang, Muyan Jiao, Kairui Wen, Yong-Jin Liu
Abstract
Fixed-capacity streaming video memory requires repeated eviction decisions whose effects accumulate over time. Yet existing policies are evaluated primarily in terms of retained information or downstream accuracy, leaving how repeated updates alter the futures supported by memory largely unexamined. We define a memory's predictive state as the future representations supported by its retained history and formulate eviction as counterfactual control over transitions in this space. We introduce SPACE (Sparse Predictive Attractor via Counterfactual Eviction), which uses a frozen multi-horizon JEPA to predict the future representations induced by alternative eviction actions. Counterfactual utility identifies future-useful alternatives, while slow predictive-basin geometry determines when to correct avoidable drift and when to adapt to sustained predictive change, without online parameter updates. We further introduce MABS-Bench, which evaluates future-task sufficiency, within-regime predictive stability, transition responsiveness, and perturbation recovery under matched causal streams and memory budgets. Across multiple video datasets, SPACE yields consistent improvements in dataset-native task performance while reducing predictive-state drift.