Legged robot walking improves with better phase aware networks

Mind the Phase: Effective Rank and Representation Health in Legged Locomotion

RoboticsArtificial Intelligence

Summary

Legged robots learn to walk and perform tricks using a type of artificial intelligence called reinforcement learning. However, it was unclear how the robot's control systems really represent the walking movements internally. The authors studied this by looking at how the robot’s control network changes during different parts of a step, called gait phases. They found that certain neural network designs better capture these phases, leading to smoother and more reliable movements when transferring from computer simulation to real robots. This insight helps make robots move more naturally and reduces jitter in their joints.

reinforcement learninglegged locomotionneural networkspolicy Jacobianeffective rankgait phaselayer normalizationresidual connectionssim-to-real transferrobot joint jitter

Authors

Felipe Tommaselli, Thiago H. Segreto, Juliano D. Negri, Ricardo V. Godoy, Marcelo Becker

Abstract

Reinforcement learning has become the leading paradigm in legged locomotion, enabling complex behaviors from backflips to parkour through massively parallel simulation. Under PPO's non-stationarity, shallow networks remain the de facto architecture, supported by carefully staged curricula and environments, yet the representations these policies learn stay poorly understood, leaving no training-time signal of how they will behave on hardware. In this work, we empirically study locomotion policies through the effective rank of the policy Jacobian and show that conditioning rank on the gait phase exposes architectural structure that global rank averages away. In particular, we find that standard architectural choices, namely layer normalization and residual connections, allocate roughly two more dimensions of effective rank to swing than to stance, which is fully absent in vanilla MLPs. Building on this, we propose a simple recipe that turns these representational signatures into smoother, more reliable sim-to-real transfer. In practice, this results in roughly 3x lower joint jitter that holds from simulation onto a physical Spot, suggesting that representation health is an effective training-time lens to track sim-to-real smoothness.