中文

具有生成模型的可证明多目标强化学习

机器学习 2021-01-12 v2 人工智能 最优化与控制

摘要

多目标强化学习(MORL)是普通单目标强化学习(RL)的扩展,适用于许多存在多个目标且相对代价未知的真实世界任务。我们研究单策略 MORL 问题,即在给定目标偏好下学习最优策略。现有方法需要强假设,例如精确已知多目标马尔可夫决策过程,并在无限数据和时间的极限下进行分析。我们提出了一种称为基于模型的包络值迭代(EVI)的新算法,它推广了 Yang 等人 2019 年的包络多目标 QQ 学习算法。我们的方法能够以多项式样本复杂度和线性收敛速度学习近最优值函数。据我们所知,这是对 MORL 算法的首个有限样本分析。

关键词

引用

@article{arxiv.2011.10134,
  title  = {Provable Multi-Objective Reinforcement Learning with Generative Models},
  author = {Dongruo Zhou and Jiahao Chen and Quanquan Gu},
  journal= {arXiv preprint arXiv:2011.10134},
  year   = {2021}
}

备注

10 pages, Workshop on Real-World Reinforcement Learning at the 34th Conference on Neural Information ProcessingSystems (NeurIPS 2020), Vancouver, Canada