Steering Generative Reinforcement Learning into Stable Robotic Controller

2026-06-15 • Robotics

Robotics

AI summaryⓘ

The authors address a problem in robot control where random action choices, though good for learning, make precise movements unstable. They introduce SteerGenPO, a method that learns a way to guide a trained action generator deterministically instead of randomly, making the robot's movements steadier. This separates the learning phase, which explores many actions randomly, from the execution phase, which controls the robot reliably. They tested their method on simulated tasks and a real robot, finding it worked better and produced more consistent behaviors than other methods.

reinforcement learningdiffusion policiesgenerative policieslatent spacedeterministic controlstochastic explorationrobot locomotionIsaac LabUnitree G1policy learning

Authors

Yixuan Wang, Shutong Ding, Ke Hu, Tianxiang Gui, Jingya Wang, Ye Shi

Abstract

Diffusion and flow-based generative policies provide a powerful policy class for reinforcement learning by inducing rich stochastic exploration through iterative action generation. However, the stochasticity of diffusion policies is not suitable for stable and precise control in high-dimensional robotic systems, where small action variations can accumulate into inconsistent motion and reduced robustness. To address this issue, we propose SteerGenPO, a latent-space reinforcement learning framework that steers a trained generative policy into a robust deterministic robotic controller. The key idea is to replace stochastic latent sampling of the trained generative policy with a learned latent actor that predicts a state-dependent latent input for the generative policies. This separates exploration and control: stochastic generative sampling provides diverse action proposals during policy learning, while deterministic latent steering provides stable and adaptive control at deployment. We evaluate SteerGenPO on six Isaac Lab benchmarks and a Unitree G1 locomotion task. The results show SteerGenPO improves over both classical RL and generative RL baselines, while its deterministic latent steering produces more stable inference-time behaviors and more reliable command responses.

View PDFOpen arXiv