Pedestrian and vehicle movements predicted more accurately and efficiently together
PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout
RoboticsArtificial IntelligenceMultiagent Systems
Summary
Predicting how pedestrians and vehicles move together is tricky because they move in very different ways. Pedestrians bend and turn with many joints, while vehicles move like rigid boxes. The authors made a model called PV-WM that predicts both pedestrian joint movements and vehicle positions in a single, combined system using only past movement information. This model is more accurate and faster than previous methods, helping computers better understand and anticipate road users' behaviors. The work could improve things like self-driving car safety and traffic management.
pedestrian articulationvehicle kinematicsworld modelroot motionjoint pose predictionrecurrent neural networkforecastingWaymo datasetaverage displacement errormulti-agent tracking
Authors
Haozhuang Chi, Jingsong Liang, Ziying Song, Lei Yang, Shihao Li, Haoruo Zhang, Chen Lv
Abstract
Local pedestrian-vehicle forecasting spans heterogeneous physical scales: pedestrians combine root locomotion with articulated motion, whereas vehicles are rigid bodies described by kinematic state and oriented extent. Existing road-agent forecasters typically omit pedestrian articulation, while pose forecasters leave vehicle futures outside the learned rollout. We introduce PV-WM, a history-only world model over structured post-perception tracks. It recurrently advances pedestrian root motion, 15-joint articulation, and learned vehicle states within a synchronized heterogeneous state. The generated pedestrian and vehicle chunks supply the next recurrent boundary; vehicle boxes are reconstructed from predicted center and heading with observed extent, and P-V geometry is recomputed after every transition. Relative to a matched one-shot complete-state predictor, recurrent execution reduces Root ADE by 12.7% and MPJPE by 14.8%. Feedback interventions show that later predictions depend on the content, temporal order, and pedestrian identity of generated articulation. Across 824 aligned Waymo contexts, with 797 providing valid future vehicle support, PV-WM reduces Root ADE by 5.2%, MPJPE by 7.6%, P-V distance error by 11.9%, and oriented-box closest-approach error by 5.8% relative to a validation-selected Modular Specialist. The single-network model uses 57.1% fewer parameters, 96.5% lower average FLOPs per local scene, and 25.5% lower measured p95 latency. PV-WM unifies this heterogeneous future state while preserving type-specific pedestrian and vehicle dynamics.