Distributional Soft Bellman Operator under the Cramér Geometry
2026-07-20 • Machine Learning
Machine Learning
AI summaryⓘ
The authors study a method called Distributional Soft Policy Iteration (DSPI) used in reinforcement learning that involves estimating distributions of outcomes rather than just averages. They focus on a specific way to measure differences between these distributions using the Cramér metric, which relates to the cumulative distribution function. They prove that the key evaluation step in DSPI reliably converges to a unique solution under this metric, given reasonable conditions on rewards and entropy. Their work also provides an equivalent understanding of the problem in a spectral (Hilbert) space, giving a solid mathematical foundation for analyzing and improving DSPI algorithms.
Distributional reinforcement learningSoft policy iterationBellman operatorCramér metricCumulative distribution functionContraction mappingEntropy regularizationHilbert spacePolicy evaluationMaximum-entropy control
Authors
Keru Wang, Yixin Deng, Yao Lyu, Stephen Redmond, Shengbo Eben Li
Abstract
Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns. Theoretical analysis of such an evaluation step requires a probability metric under which Bellman updates can be controlled, typically by showing that the operator contracts the distance between any two candidate return-distribution estimates. In this paper, we focus on the Cramér geometry, a cumulative distribution function (CDF)-based metric with an $L^2$ structure, and study whether the fixed-policy distributional soft Bellman operator has this contraction property and hence a unique fixed point under this metric. Working directly on an admissible CDF field domain, we formulate the CDF-level distributional soft Bellman operator, prove that it is a $\sqrtγ$-contraction, and obtain the corresponding unique fixed point together with convergent iterative policy evaluation. The CDF formulation also shows that this finite-Cramér-domain property follows from a uniform first-moment condition on the combined one-step reward entropy shift, rather than from separate uniform boundedness assumptions on the reward and entropy terms. We then transport the same evaluation problem to the spectral domain by conjugation, obtaining an equivalent Hilbert-space representation of the same decision process. Taken together, these results identify the Cramér-geometric Bellman fixed point associated with the policy-evaluation step of DSPI, providing a reference point for studying approximate critics, evaluation error, and critic-loss design in DSPI-style algorithms.