Joint modeling improves human and object movement prediction
Harnessing Coupled Stream Completion For Human-Object Ineraction Modeling
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Predicting how people interact with objects involves tracking body movements, hand positions, and how objects move or rotate. These parts change differently and must stay in sync to look right. The authors created a new way called TRACE that keeps each part separate but makes them update together, helping the model understand how body, hands, and objects should move in harmony. This method also allows guessing missing information if one part isn’t known and helps computers better recognize human-object interactions with language tools.
What this means in practice
- •For animation developers: Create more realistic human-object interaction animations by jointly modeling body, hand, and object movements for smoother coordinated motion.
- •For roboticists: Improve robot understanding of human-object interactions by using coupled motion completion to infer missing movement details.
Authors
Dawei Guan, Di Yang, Jiangtao Wang
Abstract
Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories & rotations, and hand articulation to remain coordinated. These components differ in scale and dynamics, but must agree on contact, relative pose, and timing. A shared representation may limit the distinct structure of each stream, while independent generation prevents each stream from responding to changes in the others. Latent supervision alone also does not directly constrain contact after decoding. We propose TRACE, a continuous latent framework that keeps stream states separate and couples their updates. TRACE encodes body, object, and hand motion into separate latents and predicts each stream velocity from the complete current interaction state. Geometric losses on decoded motion further constrain contact and object-relative motion over time. The same model supports completion of any single absent stream from the other two. Frozen flow features also serve as input to a language model for HOI understanding. Experiments on InterAct, OMOMO, and BEHAVE show that joint completion training improves generation and that frozen flow features improve understanding over raw-motion encoding. On InterAct, TRACE achieves the highest contact precision, recall, and F1 among the compared methods.