Synthetic video dataset improves accuracy of walking pattern analysis

SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation

Computer Vision and Pattern Recognition

Summary

Understanding how people walk can help doctors assess mobility issues, but current video datasets of walking are small and limited. The authors created a large synthetic video dataset called SynthGait-19K, which uses computer animations based on motion capture data to simulate many different people walking from various angles. They developed a tool called Gait2Vid to make these videos and checked that the animations matched real walking patterns. Using this dataset, they tested different methods to estimate walking features and found that training on synthetic data helps improve results on real videos. However, some walking measurements are more affected by changes in video style than others.

gait analysissynthetic datasetmotion captureSMPL modelvideo synthesisgait parametersdomain shifthuman mesh recoverypose estimation

Authors

Soroush Mehraban, Xin Lei Lin, Vida Adeli, Majid Mirmehdi, Amirhossein Dadashzadeh, Clint Hansen, Andrea Iaboni, Babak Taati

Abstract

Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.