中文

无限时域平均奖励 MDP 中基于新型策略梯度方法的阶数最优后悔界

机器学习 2025-05-13 v2

摘要

在无限时域平均奖励马尔可夫决策过程(MDP)的背景下,我们提出了两种具有一般参数化的基于策略梯度的算法。第一种算法采用隐式梯度传输进行方差缩减,确保了 O~(T2/3)\tilde{\mathcal{O}}(T^{2/3}) 量级的期望后悔界。第二种方法基于 Hessian 技术,确保了 O~(T)\tilde{\mathcal{O}}(\sqrt{T}) 量级的期望后悔界。这些结果显著改进了现有技术水平 O~(T3/4)\tilde{\mathcal{O}}(T^{3/4}) 的后悔界,并达到了理论下界。我们还证明了平均奖励函数是近似 LL-平滑的,这一结果在早期的研究中曾被假设。

关键词

引用

@article{arxiv.2404.02108,
  title  = {Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs},
  author = {Swetha Ganesh and Washim Uddin Mondal and Vaneet Aggarwal},
  journal= {arXiv preprint arXiv:2404.02108},
  year   = {2025}
}

备注

In the Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (AISTATS), 2025