StableMimic: Smooth Human-Like Recovery for Humanoid Motion Tracking - Learning Beyond the Tracking Distribution for Structured Post-Fall Behavior

2026-08-03Robotics

Robotics
AI summary

The authors developed StableMimic, a robot motion tracker that can handle situations where the robot falls and ends up in unusual positions that it hasn't seen before. Unlike normal trackers that just follow set movements and can get confused after a fall, StableMimic has special experts for both normal tracking and recovery, switching smoothly between them. It learns recovery from human get-up motions without needing explicit instructions during use, allowing the robot to stand up and continue its tasks safely. Tests showed it outperforms other methods in both tracking normal movements and recovering from falls, improving robot stability and safety.

humanoid robotmotion trackingfall recoveryreinforcement learningstate-action distributionproprioceptionget-up referencetracking policyexpert systemsrobot stability
Authors
Weihao Wu, Ming Huang, Ruofei Liu, Jinglei Nie, Shuxiang Guo, Chunying Li
Abstract
Humanoid motion trackers perform reliably within learned tracking distributions, but falls can move the robot into low-height, contact-rich states from which an advancing command is temporarily unreachable. Tracking-only policies may chase infeasible references, producing rapid, large-amplitude limb corrections that increase risk to the robot and its surroundings. We present StableMimic, a unified tracker trained beyond the nominal tracking distribution. Perturbed resets around multiple human get-up references expose prone, supine, off-balance, and intermediate ground-contact states, shaping structured recovery that returns the robot to the trackable region. Because tracking and recovery occupy markedly different state--action distributions, StableMimic uses dedicated experts for each regime and a proprioceptive gate that continuously blends their actions. A hidden successor-state objective teaches human-reference-shaped recovery without exposing reference identity or phase to the deployed Actor; deployment requires no get-up reference, recovery command, trajectory retrieval, or external policy switch. On the complete retargeted LAFAN1 dance subset, StableMimic achieves the lowest errors on all four tracking metrics among five methods. Across 100 matched push-to-fall trials per method, it recovers in 100/100 and attains the lowest values on six of seven post-fall motion and load measures, supporting improved interaction safety under this protocol. Real Unitree G1 dance and standing-reference deployments qualitatively demonstrate bounded limb motion, autonomous recovery, and command resumption.