中文

非对称演员-评论家算法的理论依据

机器学习 2025-09-09 v3 机器学习

摘要

在部分可观测环境的强化学习中,已developed许多成功算法纳入非对称学习范式。该范式利用在训练时间可获得的额外状态信息以实现更快的学习。虽然这些方法的提出学习目标通常是理论上合理的,但这些方法仍缺乏其潜在益处的精确理论依据。我们通过适应此设置中的有限时间收敛分析,为使用线性函数逼近器的非对称演员-评论家算法提出这样一种依据。 resulting finite-time bound reveals that the asymmetric critic eliminates error terms arising from aliasing in the agent state。

关键词

引用

@article{arxiv.2501.19116,
  title  = {A Theoretical Justification for Asymmetric Actor-Critic Algorithms},
  author = {Gaspard Lambrechts and Damien Ernst and Aditya Mahajan},
  journal= {arXiv preprint arXiv:2501.19116},
  year   = {2025}
}

备注

8 pages, 31 pages total