Stay Seated: Learning Omnidirectional Humanoid Locomotion on a Passive Mobile Chair with Casters
Robotics
Summary
The authors studied how a humanoid robot can move while sitting on a passive mobile chair, unlike humans who usually rely on chairs to support their weight. They created a learning setup where the robot controls the chair and moves in any direction using only internal sensors and speed commands, without sensing the chair directly. Their trained robot was able to follow movement commands well and sometimes performed better than standing. They also tested different training tricks that affected energy use and accuracy, finding combinations that improved performance. Finally, they successfully transferred their learned behavior to a real robot, showing practical use.
humanoid robotquasi-direct-drive actuatorseated locomotionpassive mobile chairreinforcement learningpolicy learningproprioceptionsim-to-real transfercost of transportvelocity tracking
Authors
Kango Yanagida, Kazuki Miyazawa, Takato Horii
Abstract
Humanoid robots with quasi-direct-drive actuators continuously generate joint torque while standing, whereas seated humans delegate weight support to chairs during desk work. As a first step toward seated loco-manipulation, we study omnidirectional seated locomotion on a passive mobile chair, requiring unfixed pelvis-seat contact and intermittent foot-floor propulsion of the robot-chair system. We extend a standard standing velocity-tracking environment with a passive-chair model, seated-state rewards, critic-only chair observations, and task-tailored contact settings. The policy is learned without motion-imitation rewards; its actor uses only proprioception and velocity commands, without contact sensing or chair states. In random-command evaluation, the policies tracked omnidirectional commands through nearly all 20-s rollouts, and the best seated policies could outperform the Standing policy in velocity tracking. Across four training seeds, a $2^3$ full-factorial comparison of symmetry regularization (SY), foot-slip regularization (FS), and command curriculum (CC) showed that FS reduced CoT but increased tracking error and that some FS-only policies converged to stationary local optima. Combining FS with either SY or CC avoided this failure without retuning FS, while SY improved bilateral leg symmetry during longitudinal motion. Direction-resolved analysis showed CoT ordered backward $<$ lateral $\ll$ forward, with planted-leg extension in backward and lateral motion and knee flexion following heel contact in forward motion. The learned policy achieved zero-shot sim-to-real transfer to a Unitree G1 and generated omnidirectional seated locomotion.