非对称演员-评论家算法的理论依据
机器学习
2025-09-09 v3 机器学习
摘要
在部分可观测环境的强化学习中,已developed许多成功算法纳入非对称学习范式。该范式利用在训练时间可获得的额外状态信息以实现更快的学习。虽然这些方法的提出学习目标通常是理论上合理的,但这些方法仍缺乏其潜在益处的精确理论依据。我们通过适应此设置中的有限时间收敛分析,为使用线性函数逼近器的非对称演员-评论家算法提出这样一种依据。 resulting finite-time bound reveals that the asymmetric critic eliminates error terms arising from aliasing in the agent state。
引用
@article{arxiv.2501.19116,
title = {A Theoretical Justification for Asymmetric Actor-Critic Algorithms},
author = {Gaspard Lambrechts and Damien Ernst and Aditya Mahajan},
journal= {arXiv preprint arXiv:2501.19116},
year = {2025}
}
备注
8 pages, 31 pages total