Stochastic control learning algorithms handle unknown dynamics and rewards
Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence
Machine Learning
Summary
Many decision-making problems involve uncertainty, and often the details about how the system behaves or what rewards are gained are unknown. The authors study a way to learn optimal decisions over time in such uncertain and complex situations using a reinforcement learning method. They design algorithms that can find the best strategies and estimate their value, even when key parts of the system are not known. Their approach includes mathematical proofs of convergence and practical tests to show the methods work.
What this means in practice
- •For financial risk managers: Improve portfolio strategies by learning optimal controls without knowing exact market drift or rewards in complex continuous-time models.
- •For robotics engineers: Develop control policies for robots operating under uncertain dynamics and rewards without explicit knowledge of system drift behavior.
Authors
Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
Abstract
We study continuous-time and possibly high-dimensional stochastic control problems where drift coefficients and running reward functions are unknown. Due to these missing model primitives, we take the exploratory, reinforcement learning (RL) framework of Wang, Zariphopoulou, and Zhou(2020) with relaxed controls and entropy regularization. The objective is to develop theoretically grounded, efficient and scalable RL algorithms to learn both the optimal value functions (which also solve the exploratory HJB equation) and optimal exploratory feedback control policies. When the diffusion coefficients do not contain control, we employ probabilistic representations of both the optimal value function and its gradient based on an auxiliary state process depending only on the diffusion part of the original dynamics. With a delicate analysis on some properly defined mappings and their fixed points, this leads to the introduction of our policy iteration algorithms and their convergence. We demonstrate the performance of our algorithms through various numerical examples. Finally, we study a special control-dependent diffusion case where probability representation of the Hessian is called for.