中文

单时间尺度带动量演员-评论家的最优样本复杂度

机器学习 2026-05-08 v2 机器学习

摘要

我们确定在有限状态-动作空间的无限时段折扣MDP中,使用单时间尺度演员-评论家算法获得ϵ\epsilon-optimal全局策略的最优样本复杂度为O(ϵ2)O(\epsilon^{-2}),超越了先前的最佳状态为O(ϵ3)O(\epsilon^{-3})。我们的ethods应用STORM(STOchastic Recursive Momentum)以减少评论家更新中的方差。然而,由于样本来自由演化策略诱导的非平稳占用度,仅用STORM进行方差缩减不足。为解决这一挑战,我们维护一小部分最近样本的缓冲区,并对每个评论家更新进行均匀抽样。重要的是,这些机制与现有的深度学习架构兼容,只需轻微修改,不会妥协实际可应用性。

关键词

引用

@article{arxiv.2602.01505,
  title  = {Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum},
  author = {Navdeep Kumar and Tehila Dahan and Lior Cohen and Ananyabrata Barua and Giorgia Ramponi and Kfir Yehuda Levy and Shie Mannor},
  journal= {arXiv preprint arXiv:2602.01505},
  year   = {2026}
}

备注

Following further internal verification, we identified foundational issues in the analytical framework, including unresolved problems in the treatment of nonstationary sampling and parts of the coupled convergence analysis under the stated assumptions. Addressing these issues requires a substantial overhaul of the theoretical framework beyond a standard revision