中文

一种用于优化平均奖赏的批量离策略Actor-Critic算法

机器学习 2016-07-19 v1 机器学习

摘要

我们开发了一种离策略Actor-Critic算法,用于从由多个个体的数据组成的训练集中学习最优策略。该算法的开发旨在应用于移动健康领域。

关键词

引用

@article{arxiv.1607.05047,
  title  = {A Batch, Off-Policy, Actor-Critic Algorithm for Optimizing the Average Reward},
  author = {S. A. Murphy and Y. Deng and E. B. Laber and H. R. Maei and R. S. Sutton and K. Witkiewitz},
  journal= {arXiv preprint arXiv:1607.05047},
  year   = {2016}
}