Papers for

animation developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Bidirectional motion and text generation improves language and movement modeling

BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion

Abstract: Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs. The project page is available at https://wengwanjiang.github.io/BiMoGen-Page.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Generating text descriptions from human motion and creating motion from text are both important but challenging tasks, especially when the two need to work together. The authors propose a new method called BiMoGen that models these tasks in both directions simultaneously, using a way to predict parts of a sequence in any order to reduce errors that accumulate with one-direction models. They also add training tricks that help the model learn better and correct its own mistakes during generation. Testing showed that BiMoGen performs well on standard motion-text datasets, improving how well motion and text correspond to each other.
Open → 2609.35407v1

Joint modeling improves human and object movement prediction

Harnessing Coupled Stream Completion For Human-Object Ineraction Modeling

Abstract: Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories & rotations, and hand articulation to remain coordinated. These components differ in scale and dynamics, but must agree on contact, relative pose, and timing. A shared representation may limit the distinct structure of each stream, while independent generation prevents each stream from responding to changes in the others. Latent supervision alone also does not directly constrain contact after decoding. We propose TRACE, a continuous latent framework that keeps stream states separate and couples their updates. TRACE encodes body, object, and hand motion into separate latents and predicts each stream velocity from the complete current interaction state. Geometric losses on decoded motion further constrain contact and object-relative motion over time. The same model supports completion of any single absent stream from the other two. Frozen flow features also serve as input to a language model for HOI understanding. Experiments on InterAct, OMOMO, and BEHAVE show that joint completion training improves generation and that frozen flow features improve understanding over raw-motion encoding. On InterAct, TRACE achieves the highest contact precision, recall, and F1 among the compared methods.

Sat 26 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Predicting how people interact with objects involves tracking body movements, hand positions, and how objects move or rotate. These parts change differently and must stay in sync to look right. The authors created a new way called TRACE that keeps each part separate but makes them update together, helping the model understand how body, hands, and objects should move in harmony. This method also allows guessing missing information if one part isn’t known and helps computers better recognize human-object interactions with language tools.
Open → 2609.32551v1

Motoneuron-inspired sampling improves model predictive path integral control smoothness

Motoneuron-Inspired Sampling for Model Predictive Path Integral Control

Abstract: Model Predictive Path Integral (MPPI) control relies on stochastic trajectory sampling, and its performance under limited rollout budgets depends strongly on the structure of the proposal distribution. Standard implementations commonly perturb control sequences with Gaussian noise, despite growing evidence that temporally correlated and structured sampling can improve finite-budget control. We introduce Spike-MPPI, a motoneuron-inspired proposal that generates temporally structured perturbations through a simplified model of motoneuron dynamics. The proposal is evaluated within a common MPPI framework on torque-actuated and antagonistically actuated MuJoCo Ant models against standard Gaussian sampling and spectrum-matched Gaussian controls. Results show that structured sampling substantially improves executed-control smoothness, while its effect on task performance depends on rollout condition and robot actuation. Spectrum matching reproduces a substantial part of the observed behavior, while the full Spike proposal retains additional effects beyond second-order spectral structure. These results support treating proposal design as a combination of second-order spectral structure and higher-order statistical organization.

Wed 23 SeptRobotics
The gist
Controlling robots by predicting their future movements often requires testing many possible actions, which can be slow and inefficient. The authors created a new way to suggest these actions by mimicking how nerve cells (motoneurons) activate muscles over time, producing smoother and more natural movement suggestions. This new method was tested on simulated robot models and showed smoother control, though the task success varied depending on the robot and conditions. Their findings suggest that both the timing and more complex features of these nerve-inspired signals help improve robot control.
Open → 2609.28325v1

Learn motor skills by combining simple building blocks in a shared network

Learning Options for Compositional Motor Control with Adapter Banks

Abstract: Learning flexible motor primitives is a hallmark of skilled motor control. Recent neuroscience theory proposes that motor primitives may be implemented as low-rank perturbations of a shared recurrent network, but leaves open how such a system is learned. We translate this principle into a novel architecture for learning motor skills end-to-end: a shared recurrent core modulated by a bank of residual adapters, each selected by a discrete latent code. Trained on closed-loop biomechanical control, the adapters develop emergent low-rank perturbations of the recurrent dynamics despite no architectural rank constraint, placing task representations in disparate subspaces of the shared core network. A simple high-level policy over the learned options, optimized while the whole network is frozen, sequences the low-rank adapters to produce novel out-of-distribution movements. We demonstrate the ability to generalize to novel motor sequences within the closed-loop control setting, improving on the generalization error of a task-input-conditioned multitask baseline by upto order of magnitude.

Tue 15 SeptMachine LearningRobotics
The gist
Controlling movements well relies on mastering reusable parts of motion. The authors created a neural system that shares a core network and quickly adapts it using small changes selected from a set of options. These options naturally organize into distinct styles of motion without forcing it. By combining these learned building blocks with a simple controller, the system can create new motor sequences it hasn't seen before, improving accuracy over traditional approaches.
Open → 2609.17042v1