Forecasting future human interactions and movements in 3D from wearable video

From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Predicting where and how people will move and interact with objects in 3D space is important for things like helping robots assist people and improving human-computer interaction. The authors created a large dataset that links future interaction locations with full body poses from first-person video. They developed a new method called HIGFlow that first predicts where interactions will happen and then uses this to guide predictions of how the body will move. Their approach better combines understanding of the environment with detailed motion over time, improving predictions compared to earlier methods.

egocentric video4D interaction forecasting3D localizationbody pose estimationmotion predictionsemantic groundingvisual dynamicsFlow Matchingassistive roboticshuman-computer interaction

Authors

Qiaohui Chu, Haoyu Zhang, Meng Liu, Haoxiang Shi, Dongmei Jiang, Liqiang Nie

Abstract

Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize them, providing an important capability for assistive robotics and human-computer interaction. Existing methods struggle to translate semantic understanding into precise continuous 3D localization and to balance motion diversity with structural consistency in pose forecasting. More fundamentally, these tasks are often modeled separately, leaving the continuous geometric and temporal correspondence between interaction locations and body motion insufficiently captured. To address these challenges, we introduce Coherent4D, a large-scale egocentric dataset for continuous 4D interaction forecasting, comprising approximately 233K samples across three domains. Each sample pairs a sequence of future 3D interaction locations with corresponding full-body poses, aligned in time and expressed in a shared coordinate system. We also provide evaluation metrics in continuous space. Building on this formulation, we propose HIGFlow, a Hand Interaction Guided Residual Flow framework that models forecasting as a cascaded where-to-how process. HIGFlow first forecasts continuous future interaction locations by combining semantic grounding with short-horizon visual dynamics, and then uses the predicted location sequence to condition a deterministic motion anchor and residual Flow Matching for diverse yet structurally consistent full-body motion forecasting. Extensive experiments across all three domains demonstrate consistent improvements over representative baselines on both location and pose forecasting, while ablations validate the contributions of the proposed components. The project page is available at https://corrineqiu.github.io/from-where-to-how/.