Papers for

video analysis engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Self supervised learning improves for continuous video streams

I Have a Stream: Making Self-Supervised Learning Work on Continuous Video

Abstract: Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.

Wed 30 SeptComputer Vision and Pattern Recognition
The gist
Training AI to understand video usually involves mixing up images from many videos and showing them repeatedly. The authors studied training AI on long, continuous video streams in the order they occur, more like how babies learn. They found that many existing learning methods struggle, mainly because the video frames in each training batch look very similar. They propose a method called StreamMAE that changes how the video is fed into the AI and selects video parts with movement, resulting in better learning on continuous videos.
Open → 2609.40333v1

Prefix-Steered memory improves long video understanding with limited frames

PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding

Abstract: Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches. We introduce Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs). PREM separates video ingestion from query answering: a recurrent writer distills visual streams into a compact 256 KiB multi-slot associative state, while a question-conditioned readout adds memory-derived key/value (K/V) steering modulations to existing non-visual prompt prefixes during prefill. This enables write-once, query-many inference without extra prompt tokens or decoding recurrence. Across six long-video benchmarks in offline and streaming end-of-stream settings, PREM consistently outperforms frozen baselines at every evaluated visual budget. Under a constrained budget of 16 frames, PREM improves macro-average accuracy by 3.06% on Qwen2.5-VL-3B, with gains of 11.0% on action antonym identification and 9.9% on localized needle retrieval. These gains require tuning 0.24% of backbone parameters at 0.03 GiB of peak GPU memory overhead.

Sun 20 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Understanding very long videos is challenging because computers can only analyze a limited number of frames at once. The authors introduce a new memory method called PREM that summarizes visual information into a small, reusable memory. This memory helps answer questions about the video without repeatedly processing all frames or adding extra data. Their method improves accuracy on several video tasks using less computational resources than previous approaches.
Open → 2609.23601v1