The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
Abstract
We study the sample complexity of the plug-in approach for learning -optimal policies in average-reward Markov decision processes (MDPs) with a generative model. The plug-in approach constructs a model estimate then computes an average-reward optimal policy in the estimated model. Despite representing arguably the simplest algorithm for this problem, the plug-in approach has never been theoretically analyzed. Unlike the more well-studied discounted MDP reduction method, the plug-in approach requires no prior problem information or parameter tuning. Our results fill this gap and address the limitations of prior approaches, as we show that the plug-in approach is optimal in several well-studied settings without using prior knowledge. Specifically it achieves the optimal diameter- and mixing-based sample complexities of and , respectively, without knowledge of the diameter or uniform mixing time . We also obtain span-based bounds for the plug-in approach, and complement them with algorithm-specific lower bounds suggesting that they are unimprovable. Our results require novel techniques for analyzing long-horizon problems which may be broadly useful and which also improve results for the discounted plug-in approach, removing effective-horizon-related sample size restrictions and obtaining the first optimal complexity bounds for the full range of sample sizes without reward perturbation.
Cite
@article{arxiv.2410.07616,
title = {The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis},
author = {Matthew Zurek and Yudong Chen},
journal= {arXiv preprint arXiv:2410.07616},
year = {2025}
}
Comments
Accepted to 36th International Conference on Algorithmic Learning Theory (ALT 2025)