Learn motor skills by combining simple building blocks in a shared network
Learning Options for Compositional Motor Control with Adapter Banks
Machine LearningRobotics
Summary
Controlling movements well relies on mastering reusable parts of motion. The authors created a neural system that shares a core network and quickly adapts it using small changes selected from a set of options. These options naturally organize into distinct styles of motion without forcing it. By combining these learned building blocks with a simple controller, the system can create new motor sequences it hasn't seen before, improving accuracy over traditional approaches.
What this means in practice
- •For robotics engineers: Sequence learned motor primitives to control robots performing new tasks without retraining the entire model.
- •For animation developers: Generate novel character animations by combining smaller learned motion segments within a shared control framework.
Authors
Sreejan Kumar, Marcelo Mattar, Lea Duncker
Abstract
Learning flexible motor primitives is a hallmark of skilled motor control. Recent neuroscience theory proposes that motor primitives may be implemented as low-rank perturbations of a shared recurrent network, but leaves open how such a system is learned. We translate this principle into a novel architecture for learning motor skills end-to-end: a shared recurrent core modulated by a bank of residual adapters, each selected by a discrete latent code. Trained on closed-loop biomechanical control, the adapters develop emergent low-rank perturbations of the recurrent dynamics despite no architectural rank constraint, placing task representations in disparate subspaces of the shared core network. A simple high-level policy over the learned options, optimized while the whole network is frozen, sequences the low-rank adapters to produce novel out-of-distribution movements. We demonstrate the ability to generalize to novel motor sequences within the closed-loop control setting, improving on the generalization error of a task-input-conditioned multitask baseline by upto order of magnitude.