Multi-token approach improves video object segmentation accuracy

MoVISA: Multi-Token Reasoning for Video Object Segmentation

Computer Vision and Pattern Recognition

Summary

Video object segmentation is about identifying and tracking objects throughout a video. Previous methods used just one text token to label objects, which sometimes made it hard to keep track of multiple or changing objects accurately. The authors designed MoVISA, a method that uses multiple tokens to represent each object across different frames, helping the system understand and locate objects better over time. This approach led to noticeable improvements on several challenging video datasets, making results more precise and easier to interpret.

What this means in practice

  • For video editing professionals: Improve software tools to accurately segment and track multiple objects in video sequences for post-production tasks.$Commercial implications: Enables development of advanced video editing features that precisely isolate multiple objects, appealing to media production companies and professionals.
  • For autonomous vehicle developers: Enhance perception systems by better identifying and tracking multiple moving objects in continuous video feeds for safer navigation.

Authors

Ruining Zhao, Ho Kei Cheng, Alexander G Schwing

Abstract

Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.