Deep value functions improve exploration in reinforcement learning
Deep Epistemic Value Functions for Optimistic Exploration
Machine Learning
Summary
Exploration is a big challenge in AI, where a program tries to learn by trying new things. The authors study why previous methods that measure uncertainty to explore often fail or behave unpredictably. They propose a new approach called DEVOTE that better manages uncertainty and keeps learning stable. Their experiments show DEVOTE finds new things and performs better on tasks than other leading methods.
What this means in practice
- •For robotics engineers: Design robots that explore complex environments more reliably by using DEVOTE’s improved uncertainty management.
- •For game ai developers: Create game agents that discover new strategies and states better through DEVOTE’s stable exploration approach.
Authors
Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause
Abstract
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.