3D point tracking improves metric accuracy using state space models

3D Point Tracking with State Space Models

Computer Vision and Pattern RecognitionMachine LearningRobotics

Summary

Tracking points in 3D space accurately is important for robots and self-driving cars, which need real-world distances rather than just pixels. The authors combine two existing tools for 2D tracking and depth estimation and add a small model to improve depth accuracy over time. This approach runs on a single common GPU and performs better than other methods using similar resources. Their method helps track points precisely in meters rather than arbitrary scales.

What this means in practice

  • For robot navigation teams: Track points in 3D accurately and efficiently on commodity GPUs for improved navigation and obstacle avoidance in robots.
  • For autonomous vehicle developers: Estimate real-world distances of tracked points from monocular cameras to inform driving decisions within hardware limits.$Commercial implications: Enables selling more accurate and efficient depth tracking systems for autonomous vehicles using only monocular cameras and standard GPUs.

Authors

Masahiro Ogawa, Qi An, Atsushi Yamashita

Abstract

Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.