Vision based model detects Parkinsons freezing of gait without wearables
Supervised Cross-Modal Feature Alignment for Zero-Wearable Freezing of Gait Detection in Parkinsonism
Computer Vision and Pattern Recognition
Summary
Freezing of gait is a movement problem in Parkinson's disease often measured by wearable sensors, but these can be inconvenient for patients. The authors developed a way to use video data instead, teaching the computer to learn from both sensor data and videos during training. This helps the video-only system recognize freezing even when parts of the body are hidden or hard to see. Their approach reaches around 85% accuracy, making it a promising step toward easier monitoring without wearing sensors.
Freezing of gaitParkinson's diseaseInertial Measurement UnitsCross-modal learningFeature alignmentVision-based detectionSkeletal trackingSubspace distillationOcclusionClassification accuracy
Authors
Aryan Singh, Chandan Biswas
Abstract
Objective assessment of Freezing of Gait (FoG) in Parkinson's disease (PD) relies predominantly on wearable Inertial Measurement Units (IMUs). While IMUs provide optimal kinematic precision, mandatory sensor attachment restricts continuous clinical deployment. Conversely, unobtrusive vision-based alternatives suffer substantial classification errors during turning-in-place tasks, where geometric self-occlusion degrades deterministic skeletal coordinates and obscures the high-frequency precursors required for FoG detection. To resolve these physical observation limits, we propose a supervised cross-modal subspace distillation framework. During optimisation, pre-trained kinematic data from IMU sensors and contextual clinical metadata act as oracles to guide a deployable visual architecture. By incorporating joint velocity and acceleration derivatives, utilising a confidence-based gating mechanism, the visual model mitigates some of the tracking errors during occlusion events. Empirical evaluations confirm this latent alignment transfers the predictive fidelity of hardware sensors directly into the visual representation, yielding $85.5\%$ accuracy, and $82.4\%$ balanced accuracy. All the while maintaining a vision only model at inference.