Humanoid robots walk safer using model-informed reinforcement learning
Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction
Robotics
Summary
Humanoid robots need to move safely in complex environments, but traditional methods can struggle with unpredictability. The authors combined a simplified physics model with modern reinforcement learning to teach robots safer walking steps. They created a safety check that adjusts robot movements in real time to prevent falls. Their approach reduced unsafe events in tests with a simulated humanoid robot, though it sometimes caused bigger side-to-side movements.
What this means in practice
- •For robotics engineers: Improve safety features in humanoid robot walking controllers by combining physics models with reinforcement learning to adjust foot placement in real time.
- •For robotics simulation teams: Evaluate and experiment with safer walking policies for humanoid robots using a physics-based simulation platform that integrates model-informed learning and safety constraints.
Authors
Victor Paredes, Ayonga Hereid
Abstract
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.