中文

应用于 Monte-Carlo AImu 实现的广义折扣函数

人工智能 2017-03-07 v1

摘要

近年来,学界在发展广义强化学习(GRL)理论方面取得了一些进展。然而,以具体方式展示这些结果的例子寥寥无几。特别是,目前没有实例展示关于广义折扣的已知结果。我们在 GRL 仿真平台 AIXIjs 中增加了为智能体分配任意折扣函数的功能,以及一个可用于确定折扣对智能体策略影响的环境。利用此平台,我们研究了几何折扣、双曲折扣和幂折扣如何影响简单 MDP 中的知情智能体。我们通过实验复现了若干理论结果,并讨论了一些相关的微妙之处。研究发现,在为 Monte-Carlo Tree Search (MCTS) 规划算法选择适当参数的前提下,智能体的行为符合理论预期。

关键词

引用

@article{arxiv.1703.01358,
  title  = {Generalised Discount Functions applied to a Monte-Carlo AImu Implementation},
  author = {Sean Lamont and John Aslanides and Jan Leike and Marcus Hutter},
  journal= {arXiv preprint arXiv:1703.01358},
  year   = {2017}
}

备注

12 pages, 4 figures