Policy gradient works best with zeroth-order reward shaping in robot control

Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

RoboticsArtificial Intelligence

Summary

Robots often learn to control themselves using rewards that guide their actions. This paper shows that for balancing and stabilizing robots, rewards based on just the robot's position (zeroth-order) are enough for learning. Adding velocity-based rewards (first-order) can actually make learning unstable and harder to tune. However, the robot's policy still needs to observe velocity to succeed. The authors’ work offers clear advice on designing rewards that make robot learning more reliable.

What this means in practice

  • For robotics engineers: Design reward functions focused on position information to improve robot balancing stability and reduce tuning effort.
  • For robotics software developers: Implement policy gradient algorithms observing velocity but shaping rewards via configuration coordinates only, to enhance learning robustness.

Authors

Yisheng Zhang, Tao Wang, Sicun Gao

Abstract

Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.