Summary
Making robots walk and move like humans on different surfaces is very hard because they need to see and adjust to the ground quickly. The authors designed a two-part system where one part plans the robot’s whole body movements from camera depth images, and the second part follows those plans exactly. They improved the planning part by letting it learn from its own experiences using a smart data search technique, which helped the robot get better at walking on new surfaces and choosing the right skills. This allows a robot dog-like humanoid to walk, run, jump on boxes, and climb stairs outdoors using just camera input without needing extra maps or sensors.
What this means in practice
- •For robotics engineers: Build multi-skill locomotion controllers for humanoid robots that adapt to varied outdoor terrain using raw camera depth images.
- •For autonomous vehicle developers: Enhance all-terrain perception and motion planning by applying multi-skill, camera-based control techniques from humanoid locomotion to off-road autonomous robots.
Abstract
General purpose humanoids require locomotion controllers that are multi-skill, perceptive, dynamic, and robust enough to go anywhere humans can. In this work, we present a two layer locomotion architecture: (1) a perceptive flow matching motion generator plans whole body trajectories from raw depth images while a (2) perceptive tracking policy trained with control-guided RL follows these motions. Both policies are trained on a library of terrain consistent motion clips created with dynamically optimized human data which yields both accurate velocity tracking and terrain consistent references. Our central contribution is a simple yet effective off-policy RL fine tuning loop that improves the motion generator. A structured search method is used with the generator to gather data for advantage weighted regression. This off-policy loop is much more sample efficient than on-policy residual fine tuning and improves terrain consistency on unseen geometries and skill compositions. We find that successful terrain traversals increased by up to 25 percentage points and skill selection improved by up to 80 percentage points. By using raw depth images to perceive the environment no odometry or height maps are needed, and outdoor deployment is easy. With two cameras, the policy can see terrain coming from further away and adjust its velocity regardless of the commanded speed so it can traverse the terrain. A single policy pair enables a Unitree G1 humanoid to walk, run, stand, jump on and off of boxes, and traverse stairs in outdoor environments. Project page: https://zolkin1.github.io/generate-track-improve/