English

Policy Gradient Learning for Distributionally Robust Markov Decision Processes under Wasserstein Ambiguity

Optimization and Control 2026-06-25 v1

Abstract

We study finite-horizon Markov Decision Processes (MDPs) under distributional uncertainty in the transition kernels and develop a policy-gradient framework for Wasserstein distributionally robust control. Ambiguity is modeled by state-action dependent Wasserstein balls around nominal transition kernels, leading to a max-min control problem over randomized policies and admissible transition laws. Since the worst-case transition law depends implicitly on the policy parameters, the usual policy-gradient argument does not apply. We address this difficulty by using a Wasserstein dual reformulation of the robust Bellman recursion and analyzing its directional differentiability. This yields an explicit recursive characterization of the robust policy gradient. Building on this characterization, we propose a robust actor-critic algorithm and illustrate its behavior on discrete and continuous benchmark examples.

Cite

@article{arxiv.2606.27610,
  title  = {Policy Gradient Learning for Distributionally Robust Markov Decision Processes under Wasserstein Ambiguity},
  author = {Yadh Hafsi and Samy Mekkaoui and Huyên Pham and Kaixin Yan},
  journal= {arXiv preprint arXiv:2606.27610},
  year   = {2026}
}
R2 v1 2026-07-22T20:10:54.931Z