Summary
Robots can learn to do tasks by copying human demonstrations, but these demonstrations often use fixed settings that are too rigid. This work shows how to turn these demonstrations into controllers that smoothly change their stiffness and resistance, making the robot better at handling contact tasks. The authors created methods to better measure and adapt the robot’s behavior during learning, resulting in smoother motions with less force fluctuation. They tested this approach on real robots doing contact-rich tasks and found improved performance and faster computation. This method also helps teach robots with clearer guidance, though some tasks still show uneven success.
What this means in practice
- •For robotics engineers: Improve robot control for delicate contact tasks by generating smooth variable-impedance controllers from fixed demonstration data.
- •For industrial automation teams: Deploy robots that learn from demonstrations with better force control and more reliable motions in manufacturing environments.
Authors
Jiahao Liu, Kento Kawaharazuka, Tasuku Makabe, Kei Okada
Abstract
CMDIR extends Manifold-Decomposed Impedance Retargeting (MDIR) to transform fixed-impedance demonstrations into continuous variable-impedance controllers, which can also serve as structured supervision for imitation learning. Continuous Task-Manifold Impedance Representation (TMIR) pairs an evolving task frame with controller instructions. Demo-relative Compromise dynamics retain moving-basis transport and control/physical metric mismatch, yielding displacement, reaction-impulse, and perturbation-sensitivity criteria. Quality-to-Fast automatically compiles a solver structure from development paths within a predefined finite space, re-instantiates that structure for each demonstration, and certifies the resulting candidate by multi-resolution evaluation. Across 225 retargeted-controller trials in three real contact tasks, full CMDIR improves mean task-proxy retention and reduces mean pose deviation, force fluctuation, and peak force relative to discrete MDIR. FastMPO achieves a $5.8$--$9.4\times$ speedup over C-MPO with comparable closed-loop outcomes. Downstream experiments demonstrate learnability of the complete TMIR supervision interface; lower force fluctuation and peak force are observed among successful executions, while completion reliability remains uneven across tasks and environments.