English

Finding the Near Optimal Policy via Adaptive Reduced Regularization in MDPs

Machine Learning 2020-11-03 v1 Machine Learning

Abstract

Regularized MDPs serve as a smooth version of original MDPs. However, biased optimal policy always exists for regularized MDPs. Instead of making the coefficient{\lambda}of regularized term sufficiently small, we propose an adaptive reduction scheme for {\lambda} to approximate optimal policy of the original MDP. It is shown that the iteration complexity for obtaining an{\epsilon}-optimal policy could be reduced in comparison with setting sufficiently small{\lambda}. In addition, there exists strong duality connection between the reduction method and solving the original MDP directly, from which we can derive more adaptive reduction method for certain algorithms.

Keywords

Cite

@article{arxiv.2011.00213,
  title  = {Finding the Near Optimal Policy via Adaptive Reduced Regularization in MDPs},
  author = {Wenhao Yang and Xiang Li and Guangzeng Xie and Zhihua Zhang},
  journal= {arXiv preprint arXiv:2011.00213},
  year   = {2020}
}
R2 v1 2026-06-23T19:48:10.085Z