Long-term video segmentation improves tracking in complex videos
LVMT: Video Mask Transformer for Long-term Video Segmentation
Computer Vision and Pattern Recognition
Summary
Tracking objects over long videos with many interruptions is difficult for existing methods. The authors suggest that this is because current systems can’t smartly choose what information to remember and struggle to learn from long videos due to memory limits. They created a new way to remember and pass object details using a special memory module and a training trick that breaks videos into smaller parts. This makes their model both more accurate and faster at segmenting objects over long videos than earlier methods.
What this means in practice
- •For video editing professionals: Segment objects accurately over long video clips even with occlusions for streamlined post-production workflows.
- •For security monitoring teams: Track objects through long surveillance footage where occlusions and complexities are common to improve incident analysis.
Authors
Narges Norouzi, Niccol`o Cavagnero, Idil Esen Zulfikar, Bastian Leibe, Gijs Dubbelman, Daan de Geus
Abstract
Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradients. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time. Second, to allow training on long videos, we introduce Truncated Query Propagation (TQP), a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients. The resulting model is called the Long-term Video Mask Transformer (LVMT). Extensive experiments on six benchmarks show that LVMT sets a new state of the art across a range of video segmentation tasks, while retaining the speed of the highly efficient model it is based on, making it 10X faster than the prior state of the art. Code: https://www.tue-mps.org/lvmt