Video segmentation predicts object movement without extra inputs
DynEoMT: Learning Object Dynamicity from Online Segmentation Queries
Computer Vision and Pattern RecognitionRobotics
Summary
Video segmentation helps find and follow objects in videos, but it can't tell if an object is moving on its own or if the whole camera is moving. The authors created DynEoMT, a method that can label each object as either moving independently or static, using only the current video frame and some previous information from earlier frames. They also developed a way to train the system without needing special labels showing which objects move. Their method works well on several video datasets while keeping good object tracking accuracy. This means it can figure out object movement online without needing complex extra data like optical flow or depth.
What this means in practice
- •For robotics engineers: Improve robots’ perception by identifying independently moving objects using only current video frames and propagated queries without expensive motion sensors.
- •For video surveillance teams: Detect moving objects more accurately in live video feeds without relying on additional sensor data like depth or camera pose.
Authors
Calvin Galagain, Martyna Poreba, François Goulette
Abstract
Video segmentation models recognize and track objects over time, but they do not indicate whether each segmented region moves independently of the observing camera. This dynamicity attribute cannot be inferred from semantics alone and is confounded by camera ego-motion. We introduce \method, an online framework that augments query-based video segmentation with region-level dynamicity prediction. It jointly produces the original segmentation outputs and a dynamic or static state for each predicted region. At inference, DynEoMT uses only the current frame and propagated queries, without optical flow, depth, camera pose, previous RGB frames, or feature maps. Because established video segmentation benchmarks do not annotate this attribute, we also introduce a class-agnostic offline supervision pipeline using camera-compensated optical flow and confidence-aware temporal filtering. Across VIPSeg, OVIS, YouTube-VIS 2022, and VSPW, DynEoMT achieves balanced accuracies of 84.3, 68.0, 68.6, and 87.6, respectively, while largely preserving segmentation performance. These results show that segmentation-region dynamicity can be learned from propagated queries, enabling its online prediction without a dedicated motion-processing pipeline at inference. The complete code will be released as open source to enable full reproduction of the method and experiments.