Papers for

motion recognition developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Motion language evaluators struggle to recognize structure and mirror actions

Can Motion-Language Models Ground Structure? STRIDE for Evaluating the Evaluators

Abstract: Motion-language models are typically scored by motion-language evaluators, but how well these evaluators ground language structure remains unclear. Here, we introduce the Structure grounding via Temporal-order, Reflection, and Identity Diagnostic Evaluation (STRIDE) benchmark to systematically evaluate the ability of evaluators to track temporal order, mirror reflection, and action identity. STRIDE comprises $5{,}869$ triples, each consisting of a motion, its original caption, and a perturbed caption, spanning both short and long descriptions. We likelihood-balance caption pairs to reduce text-only bias and estimate each evaluator's caption preference under unrelated motions to measure the discrimination gain from matched motions relative to this baseline. Our experiments reveal weak structural grounding and severe deficits in mirror sensitivity among the audited evaluators, which commonly used evaluation protocols fail to expose. To understand why these limitations go undetected in standard tests, we examine the evaluators more closely. We find that text-only priors alone can solve naive perturbation tests on existing datasets, while common retrieval and distributional metrics barely respond to structural corruption introduced by mirroring ground-truth motions. These findings suggest a natural intervention: structural hard negatives. Our experiments show that a simple modification to contrastive learning substantially improves performance on temporal order and mirror reflection. The benchmark and code will be released.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Motion-language models are tools that match actions with descriptions, but the systems used to judge how well they do this often miss important details. The authors created a new test called STRIDE that checks if these judges can tell the right order of movements, recognize when an action is mirrored, and identify the action itself. They found that these evaluators often fail to notice mirrored actions and don't properly understand the structure of the descriptions. They also discovered that many current tests can be passed without truly understanding the action, and showed a way to improve models by training them with harder examples.
Open → 2609.32462v1

Skeleton model learns motions across different sensors with one system

One for All: Generalist Foundation Model for Cross-Sensor Skeleton Representation Learning

Abstract: For learning generalizable motion representations from large-scale unlabeled data, Self-supervised learning (SSL) has become a widely adopted methodology. However, existing approaches are primarily limited by the inherent heterogeneity of skeleton data---characterized by varying joint counts, indexing protocols, and topological structures across different sensors---which typically necessitates training separate, sensor-specific, or even entirely dataset-specific models. To overcome this, we introduce SOfA (Skeleton One for All), the first generalist foundation model designed to achieve sensor-unified skeleton representation learning across diverse sensors. To accommodate the dimensional gap caused by varying joint counts, we introduce a fixed-size set of learnable Canonical Joint Slots, acting as a universal vessel that seamlessly accommodates arbitrary skeletal topologies. SOfA fills these slots via an attention mechanism that dynamically aggregates skeletal information from sensor-specific inputs. Furthermore, we resolve joint index misalignment between various sensors by introducing a Semantic Joint Embedding derived from a pre-trained text encoder, rather than relying on absolute positional embeddings. To validate our approach, we standardized ten 3D skeleton datasets for unified training. Extensive experiments demonstrate that SOfA can serve as a truly universal encoder, achieving state-of-the-art (SOTA) performance across a wide range of downstream tasks and sensor types, often outperforming dataset-specific specialist models with a single foundation model.

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
Different motion sensors record body movements in varying ways, making it hard to create one model that works for all. The authors created SOfA, a general model that can understand body movements from multiple types of sensors by using special slots for joints and clever ways to match body parts even if their labels differ. They trained this model on ten different datasets and found it works better than models made just for one sensor type. This means SOfA can learn general motion patterns that apply widely.
Open → 2609.07078v1