一种用于优化平均奖赏的批量离策略Actor-Critic算法
机器学习
2016-07-19 v1 机器学习
摘要
我们开发了一种离策略Actor-Critic算法,用于从由多个个体的数据组成的训练集中学习最优策略。该算法的开发旨在应用于移动健康领域。
引用
@article{arxiv.1607.05047,
title = {A Batch, Off-Policy, Actor-Critic Algorithm for Optimizing the Average Reward},
author = {S. A. Murphy and Y. Deng and E. B. Laber and H. R. Maei and R. S. Sutton and K. Witkiewitz},
journal= {arXiv preprint arXiv:1607.05047},
year = {2016}
}