Papers for

video processing engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Octree method speeds video processing with fewer data points

Octree-based Video Representation

Abstract: Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Videos have lots of repeated and simple parts mixed with detailed sections, but current methods treat all parts the same, making processing inefficient. The authors introduce OctVideo, a way to represent video using a tree-like structure that breaks up space and time so simple areas are stored coarsely and detailed parts finely. This method uses fewer computations, encodes and decodes faster, and still produces good quality video reconstructions. It also works well on different video datasets without extra training and helps recognize video content using fewer inputs.
Open → 2609.33100v1

Video sparse attention speeds up generation with fine-grained routing

Improving Video Sparse Attention with Fine-grained Router and Sparse Rebasing

Abstract: We present VSA2, a frontier trainable sparse attention for video DiTs. VSA2 includes a variety of new architectural features and training procedures that we apply across all stages of the DiT development cycle, including pretraining, RL, and inference, to produce a DiT with comparable or better quality than a full attention counterpart. Architecturally, VSA2 introduces a fine-grained router that improves the precision of identifying critical tokens and supports dynamic computation by allowing each query to attend to a variable number of key-value pairs. In training, we identify a Hard-to-Easy Curriculum, where models trained under high sparsity and later evaluated with lower sparsity during inference not only generalize effectively, but also outperform models trained with full attention in motion quality. VSA2 is also flexible: it can replace full attention during the middle of progressive low-to-high resolution pretraining, rebasing early-stage full-attention checkpoints. Experiments show that VSA2 reduces attention computation by half over VSA with lower loss. On 720p videos, it accelerates attention by 8.9x and end-to-end generation by 4.62x compared to the FlashAttention-3 baseline, while achieving comparable or better video quality.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Processing videos efficiently is challenging because it requires lots of computations to focus on important parts. The authors developed VSA2, a method that smartly picks which parts of the video to pay attention to, using a ‘fine-grained router’ that adjusts how much is processed for each part. They also found a special training approach that helps models perform even better when they focus less, making video generation faster without losing quality. Their method can be used during different training stages and speeds up video creation significantly compared to previous techniques.
Open → 2609.32882v1

Space improves video memory by predicting future use patterns

SPACE: Sparse Predictive Attractor via Counterfactual Eviction for Streaming Video Memory

Abstract: Fixed-capacity streaming video memory requires repeated eviction decisions whose effects accumulate over time. Yet existing policies are evaluated primarily in terms of retained information or downstream accuracy, leaving how repeated updates alter the futures supported by memory largely unexamined. We define a memory's predictive state as the future representations supported by its retained history and formulate eviction as counterfactual control over transitions in this space. We introduce SPACE (Sparse Predictive Attractor via Counterfactual Eviction), which uses a frozen multi-horizon JEPA to predict the future representations induced by alternative eviction actions. Counterfactual utility identifies future-useful alternatives, while slow predictive-basin geometry determines when to correct avoidable drift and when to adapt to sustained predictive change, without online parameter updates. We further introduce MABS-Bench, which evaluates future-task sufficiency, within-regime predictive stability, transition responsiveness, and perturbation recovery under matched causal streams and memory budgets. Across multiple video datasets, SPACE yields consistent improvements in dataset-native task performance while reducing predictive-state drift.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
When computers watch videos and remember what they see over time, they have limited memory and need to decide what to keep or forget. The authors propose a way to predict how different choices about what to remember affect future video understanding. Their method, called SPACE, uses a special model to foresee future video parts that will be useful and makes smarter decisions about what to keep in memory. This helps keep important information longer and improves overall video performance without needing to relearn during use.
Open → 2609.32592v1

KeyRec reduces memory use for understanding long videos and streams

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

Abstract: Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose \textbf{KeyRec}, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.

Sat 26 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Videos and streams are long, and computers often slow down trying to watch all the details at once. The authors designed KeyRec, a system that keeps important recent moments in short-term memory and groups earlier events in a smart way to save space. It decides how much attention to give to recent versus older parts based on the question asked, without rewatching old video parts. This makes video understanding faster and more efficient without losing important details.
Open → 2609.32182v1

Dual-stream method improves hdr video with alternating exposures

Double-stream registration with pyramid fusion for HDR video with alternating exposures

Abstract: High dynamic range (HDR) video reconstruction from al\-ter\-na\-ting-exposure sequences remains challenging, especially in regions with extreme luminance variation. We propose a novel HDR reconstruction framework based on dual-stream registration and accurate pyramid fusion. Given three consecutive frames, our method computes optical flow directly with the central frame, while introducing a complementary midpoint displacement strategy to handle cases with severe overexposition. A pyramid fusion stage then merges the resulting radiance and LDR images into a final HDR output. Experimental results demonstrate that our approach consistently outperforms state-of-the-art methods.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Capturing videos with very bright and very dark areas is hard because normal cameras can’t see all lighting details at once. The authors propose a new way to combine video frames taken with different brightness settings to create clearer, high dynamic range (HDR) videos. Their method aligns video frames more accurately and merges the information using a pyramid fusion technique, handling even very bright spots well. Tests show their approach works better than existing methods.
Open → 2609.31108v1