中文

基于Actor-Critic方法的平均奖励强化学习全局收敛性 sharper 分析

机器学习 2025-05-07 v3

摘要

本 work 检查具有通用策略参数化的平均奖励强化学习。现有的状态-of-the-art(SOTA)保证要么次优,要受到多项挑战的限制,包括:对于状态-动作空间大小的差异不佳、高迭代复杂度以及对混合时间和到达时间的依赖。为此,本文提出了基于多级蒙特卡洛的自然演员-评论家(MLMC-NAC)算法。我们的工作首次实现了对于平均奖励马尔可夫决策过程(MDPs)的全局收敛率为O~(1/T){\tilde{\mathcal{O}}}(1/\sqrt{T})(其中TT为时域长度),而无需混合时间和到达时间的知识。此外,收敛率与状态空间大小无关,因此即便适用于无限状态空间。

关键词

引用

@article{arxiv.2407.18878,
  title  = {A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic Approach},
  author = {Swetha Ganesh and Washim Uddin Mondal and Vaneet Aggarwal},
  journal= {arXiv preprint arXiv:2407.18878},
  year   = {2025}
}

备注

Accepted to the 42nd International Conference on Machine Learning (ICML), 2025