折扣马尔可夫决策过程中近极小极大最优的强化学习
机器学习
2022-01-04 v3 最优化与控制
机器学习
摘要
我们研究表格设定下折扣马尔可夫决策过程(MDPs)的强化学习问题。我们提出一种名为 UCBVI- 的基于模型的算法,该算法基于不确定性下的乐观原则(optimism in the face of uncertainty principle)与 Bernstein 型奖励项。我们证明 UCBVI- 达到了 的悔值(regret),其中 为状态数, 为动作数, 为折扣因子, 为步数。此外,我们构造了一类困难 MDPs,并证明对于任意算法,其期望悔值至少为 。我们的上界在对数因子内与极小极大下界相匹配,这表明 UCBVI- 对于折扣 MDPs 是近极小极大最优的。
引用
@article{arxiv.2010.00587,
title = {Nearly Minimax Optimal Reinforcement Learning for Discounted MDPs},
author = {Jiafan He and Dongruo Zhou and Quanquan Gu},
journal= {arXiv preprint arXiv:2010.00587},
year = {2022}
}
备注
33 pages, 1 figure, 1 table. In NeurIPS 2021