Video models keep object identities better under occlusion and overlaps
Learning to Reason with Persistent Object States for Video Instance Segmentation
Computer Vision and Pattern Recognition
Summary
Tracking objects in videos can be tricky when they get hidden or look very similar to others. The authors introduce POSReasoner, a tool that carefully checks when to update the memory of an object's position and identity, avoiding mistakes that happen from confusing objects over time. It uses a system that proposes possible matches and then verifies them before confirming updates, helping to keep track of objects accurately even after they disappear and reappear. This approach works with different existing video segmentation models and improves their accuracy, especially in difficult cases.
What this means in practice
- •For security camera operators: Improve tracking of people or objects that get hidden or reappear in surveillance videos to reduce mistaken identity switches.
- •For video editing software developers: Enable more reliable object tracking in video editing tools for effects that need consistent identification of objects across frames despite occlusions.
Authors
Yongxue Xu, Boxue Yang, Ziqian Liu, Shaoqiu Zhang, Rui Qian, Haopeng Chen
Abstract
Video segmentation models maintain object identities by carrying instance information across frames. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an unreliable update can overwrite a valid history and cause persistent identity drift. We introduce POSReasoner, a trainable, plug-and-play framework that explicitly decides when an observation should change an object's state. Each persistent state records identity, confidence, and absence history. A sparse state-observation graph supports Propose-Verify reasoning: provisional associations are revisited using object history, predicted presence, and competition among identities. The verified decisions determine whether to retain, update, reactivate, or suppress each state, while a learned gate controls the evidence written back to memory. Only verified transitions update the persistent state used in subsequent frames. POSReasoner uses standard video annotations and keeps the base model frozen, enabling integration with diverse VOS and VIS architectures. Experiments across long-term VOS and VIS benchmarks show consistent improvements over strong baselines, with the largest gains under occlusion and object reappearance.