无限时域平均奖励 MDP 中基于新型策略梯度方法的阶数最优后悔界
机器学习
2025-05-13 v2
摘要
在无限时域平均奖励马尔可夫决策过程(MDP)的背景下,我们提出了两种具有一般参数化的基于策略梯度的算法。第一种算法采用隐式梯度传输进行方差缩减,确保了 量级的期望后悔界。第二种方法基于 Hessian 技术,确保了 量级的期望后悔界。这些结果显著改进了现有技术水平 的后悔界,并达到了理论下界。我们还证明了平均奖励函数是近似 -平滑的,这一结果在早期的研究中曾被假设。
引用
@article{arxiv.2404.02108,
title = {Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs},
author = {Swetha Ganesh and Washim Uddin Mondal and Vaneet Aggarwal},
journal= {arXiv preprint arXiv:2404.02108},
year = {2025}
}
备注
In the Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (AISTATS), 2025