Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes
2026-08-24 • Machine Learning
Machine Learning
AI summaryⓘ
The authors study a new way to model decision-making where several possible actions and their outcomes are considered together, preserving how these outcomes are related under the same random events. They build on an existing framework called Joint Markov decision processes (JMDPs) to develop methods for finding the best strategies. Their work defines a new mathematical operator and shows it can be used to reliably find optimal strategies under certain conditions. They also provide ways to approximate these solutions using neural networks.
Markov decision processJoint Markov decision processcoupled dynamicsBellman operatorWasserstein distanceoptimal controldistributional reinforcement learningmoment evaluationneural approximationcounterfactual outcomes
Authors
Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi, Vijay Gupta, Abolfazl Hashemi
Abstract
Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism allows one to reason about the marginal law of each action but discards dependence across these counterfactual outcomes. The Joint Markov decision process (JMDP) formalism preserves that dependence. Prior work established the formalism and solved the fixed-policy joint moment evaluation problem in JMDPs. This paper develops optimal-control methods. We define a nonparametric distributional Bellman optimality operator for JMDPs, and prove that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law. For the first two moments, we establish convergence under a weaker condition that permits several mean-optimal actions as long as their tie resolutions share a second-moment fixed point. We also derive sampled targets for neural approximation.