English

Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs

Machine Learning 2025-05-13 v2

Abstract

We present two Policy Gradient-based algorithms with general parametrization in the context of infinite-horizon average reward Markov Decision Process (MDP). The first one employs Implicit Gradient Transport for variance reduction, ensuring an expected regret of the order O~(T2/3)\tilde{\mathcal{O}}(T^{2/3}). The second approach, rooted in Hessian-based techniques, ensures an expected regret of the order O~(T)\tilde{\mathcal{O}}(\sqrt{T}). These results significantly improve the state-of-the-art O~(T3/4)\tilde{\mathcal{O}}(T^{3/4}) regret and achieve the theoretical lower bound. We also show that the average-reward function is approximately LL-smooth, a result that was previously assumed in earlier works.

Keywords

Cite

@article{arxiv.2404.02108,
  title  = {Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs},
  author = {Swetha Ganesh and Washim Uddin Mondal and Vaneet Aggarwal},
  journal= {arXiv preprint arXiv:2404.02108},
  year   = {2025}
}

Comments

In the Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (AISTATS), 2025