DiMoP improves skeleton action recognition by learning diverse motions
DiMoP: Diffusion-Driven Motion Representation Learning With Frame-Level Pseudo-Classification for Skeleton-Based Action Recognition
Computer Vision and Pattern Recognition
Summary
Recognizing human actions from skeleton data is harder when the motions are subtle or moderate instead of strong. The authors propose DiMoP, a method that learns to represent different strengths of motion by adding and removing noise on parts of the skeleton over time. They also create a way to label frames automatically to make the learning more consistent. This approach helps computers better understand a wide range of actions and improves recognition accuracy on multiple datasets.
What this means in practice
- •For mobile app developers: Build more accurate human activity recognition features that handle subtle and varying motions reliably in smartphones and wearable devices.
- •For security system engineers: Enhance video surveillance systems with improved recognition of diverse human actions from skeleton data for better anomaly detection.
Authors
Shanaka Ramesh Gunasekara, Wanqing Li, Nikalal Kaldera, Philip Ogunbona, Jack Yang
Abstract
Robust skeleton-based action recognition requires representations that capture a wide spectrum of motions, from subtle to moderate and strong ones. Existing methods often focus on strong motions. This paper introduces DiMoP, a masking- and diffusion-driven motion representation learning method with frame-level pseudo-classification to explicitly learn the distribution of joint motions rather than regressing deterministic coordinates, as existing methods often do. By diffusing masked joints with progressive noise and denoising them conditioned on visible joints, DiMoP learns through controllable noising and denoising processes, enabling uniform learning of weak, moderate, and strong dynamics. To enable the masking-based generative diffusion learning with a discriminative capability, a pseudo-frame classifier is proposed that enforces the learning towards sequence-consistent and temporally coherent pseudo-labels without manual annotations. Together, these strategies provide a principled mechanism for joint generative and discriminative motion modeling. DiMoP achieves state-of-the-art performance across NTU RGB+D 60/120, and PKUMMD, including a 1.1 percentage point gain over prior works on NTU RGB+D 120 with the cross-subject protocol.