Papers for

video editing professionals

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Long-term video segmentation improves tracking in complex videos

LVMT: Video Mask Transformer for Long-term Video Segmentation

Abstract: Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradients. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time. Second, to allow training on long videos, we introduce Truncated Query Propagation (TQP), a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients. The resulting model is called the Long-term Video Mask Transformer (LVMT). Extensive experiments on six benchmarks show that LVMT sets a new state of the art across a range of video segmentation tasks, while retaining the speed of the highly efficient model it is based on, making it 10X faster than the prior state of the art. Code: https://www.tue-mps.org/lvmt

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Tracking objects over long videos with many interruptions is difficult for existing methods. The authors suggest that this is because current systems can’t smartly choose what information to remember and struggle to learn from long videos due to memory limits. They created a new way to remember and pass object details using a special memory module and a training trick that breaks videos into smaller parts. This makes their model both more accurate and faster at segmenting objects over long videos than earlier methods.
Open → 2609.34895v1

Video super-resolution improves low-quality videos using local frame context

LoCoVSR: Local Context Diffusion Posterior Sampling for Video Super-Resolution

Abstract: Video super-resolution (VSR) is an ill-posed inverse problem that aims to reconstruct a high-resolution (HR) video from a noisy, low-resolution (LR) version of it. We present LoCoVSR, a diffusion-based VSR framework that leverages pixel-space denoising diffusion probabilistic models. LoCoVSR integrates the Diffusion Posterior Sampling technique with spatio-temporal context learning, operating in a moving-average form. A localized window of adjacent LR frames is used for recovering each center frame, while applying a shared noise trajectory across all frames. The localized windowing enables processing of long videos without length limitations, supports parallel inference, and prevents error accumulation that may occur in recursive processing. Unlike prior methods, LoCoVSR offers a simple yet very effective VSR solution, avoiding explicit optical flow estimation, or information loss caused by latent space processing. Trained on the VFHQ face dataset, LoCoVSR achieves accurate, temporally consistent and high-quality upscaling with competitive results against recent diffusion-based VSR approaches.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Video super-resolution means turning low-quality videos into sharper, clearer high-quality ones, but it’s a tricky problem with many possible solutions. The authors created LoCoVSR, a new method that looks at small groups of nearby low-resolution video frames to better guess the details in the high-resolution video. Their method avoids complicated steps like estimating motion between frames and works well on long videos without losing quality over time. They tested LoCoVSR on face video data and showed it makes videos look sharper and smoother over time compared to similar recent methods.
Open → 2609.32742v1

Multi-token approach improves video object segmentation accuracy

MoVISA: Multi-Token Reasoning for Video Object Segmentation

Abstract: Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.

Thu 24 SeptComputer Vision and Pattern Recognition
The gist
Video object segmentation is about identifying and tracking objects throughout a video. Previous methods used just one text token to label objects, which sometimes made it hard to keep track of multiple or changing objects accurately. The authors designed MoVISA, a method that uses multiple tokens to represent each object across different frames, helping the system understand and locate objects better over time. This approach led to noticeable improvements on several challenging video datasets, making results more precise and easier to interpret.
Open → 2609.28956v1