Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking

2026-08-10Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors tackle the problem of tracking points and reconstructing 3D shapes over time more smoothly and accurately. They create a system called Uni4R that predicts how points move continuously, not just at fixed times, by combining ideas from Optimal Transport and solving equations that describe motion. Their approach uses a special decoder that learns global and local motion patterns, allowing predictions at any moment in time. They also introduce a training method that works without needing exact velocity data at every frame. Tests show their method improves over previous ones in both tracking and 4D reconstruction.

4D reconstructionpoint trackingOptimal TransportOrdinary Differential Equationvelocity fieldFlow Matchingkinematic priorintegral-consistency trainingcontinuous time modeling
Authors
Liying Yang, Hao Mo, Jialun Liu, Chen Liu, Xinxing Yu, Chenhao Guan, Hui Ma, Xiao Cao, Ajian Liu, Yanyan Liang
Abstract
Existing unified 4D reconstruction and point tracking approaches typically rely on heuristic interpolations or just predict at integer timestamps, lacking kinematic coherence and failing to model dynamics at any arbitrary timestamp. In this paper, we propose Uni4R, a framework that unifies these tasks by learning continuous velocity fields through the synergy of Optimal Transport (OT) and Ordinary Differential Equation (ODE). Importantly, this continuous velocity field acts as a kinematic prior that mutually benefits both 4D reconstruction and point tracking. Specifically, we propose the Flow Matching Guided Decoder (FMGD). A global velocity branch first extracts anchor features that capture the global dynamic state of the sequence. Then, FMGD leverages Flow Matching (FM) theory to formulate a probability path defined by OT on the anchor feature manifold, instantiating it as FM-guided velocity features for velocity prediction. This establishes a robust kinematic inductive bias. Meanwhile, a point reconstruction branch provides geometric features. The local velocity prediction module then joint above features and time embeddings, to decode velocities at arbitrary timestamps. To overcome the absence of high-quality ground-truth velocities in fractional frames, we propose an integral-consistency training strategy. This strategy uses an ODE solver to integrate velocities to recover target pointmaps, enabling the model to be supervised end-to-end directly from integer timestamps. Experimental results demonstrate that Uni4R achieves SOTA performance in both 4D reconstruction and point tracking, and achieves SOTA in our new kinematics-aware benchmark at continuous time.