Policy Gradient Learning for Distributionally Robust Markov Decision Processes under Wasserstein Ambiguity

We study finite-horizon Markov decision processes under distributional uncertainty in the transition kernels and develop a policy-gradient framework for Wasserstein distributionally robust control. Ambiguity is modeled by Wasserstein balls of common radius centered at state--action-dependent nominal transition kernels, leading to a max--min problem over randomized policies and admissible transition laws. Because the worst-case transition law depends implicitly on the policy parameters, the standard policy-gradient argument does not apply directly. We address this difficulty by combining the dynamic programming recursion with Wasserstein duality and a primal envelope argument. In general, the right and left directional derivatives of the one-step worst-case value are obtained by taking the minimum or maximum expected downstream value derivative over the set of worst-case transition laws. In finite state--action spaces, this set is characterized through the optimal face of a transport linear program, yielding an exact directional-derivative recursion. Under the required stability conditions and uniqueness of the active dual and transport optimizers, the derivative becomes linear in the policy perturbation and admits an explicit vector valued policy-gradient recursion. Building on this representation, we propose a robust actor--critic implementation and evaluate it on benchmark examples.

Paper

Similar papers

© 2026 NYSGPT2525 LLC