Free-Energy-Gated Plasticity for Real-Time Online Motor Learning in Physical Human--Robot Interaction

2026-08-24Robotics

RoboticsHuman-Computer Interaction
AI summary

The authors developed a way for robots to learn new movements while still remembering old ones during real-time interaction. They improved a special kind of neural network (PV-RNN) by adding Free-Energy-Gated Plasticity (FEGP), which adjusts how fast the robot learns based on how well it predicts its environment. In tests, this method helped the robot learn three distinct movement patterns from scratch without any prior training or breaks between tasks. The authors found that the timing of learning adjustments is more important than just how much learning happens overall. This helps the robot keep older skills while still adapting to new ones continuously.

online learningsynaptic adaptationPredictive CodingVariational Recurrent Neural Network (PV-RNN)Free Energyplasticitymotor patternshuman-robot interactionlearning ratemodel-environment mismatch
Authors
Hiroki Sawada, Jun Tani
Abstract
Fully online embodied learning requires synaptic adaptation to acquire new behaviors while preserving previously learned dynamics during ongoing interaction. We extend the Predictive-Coding-inspired Variational Recurrent Neural Network (PV-RNN) to continuously adapt its synaptic weights and propose Free-Energy-Gated Plasticity (FEGP), which regulates the effective learning rate according to variational free energy. In real-time physical human--robot interaction, a randomly initialized network acquired three cyclic motor patterns without offline pretraining, replay, or task-boundary signals, with all three patterns emerging in autonomous rollouts. Controlled experiments over ten randomized teaching streams and five network initializations per stream showed that FEGP substantially improved repertoire coverage and retention of previously acquired patterns after they left the recent observation window. Neither a constant learning rate matched to the gate's time-averaged effective rate nor replay of the same gain values with disrupted temporal organization reproduced these improvements. These results indicate that the temporal allocation of plasticity relative to model--environment mismatch, rather than simply its average magnitude or distribution, is critical for maintaining previously acquired behaviors during continued online learning.