用于星际争霸 II 的异步优势 actor-critic 智能体
人工智能
2018-07-25 v1
摘要
深度强化学习,尤其是异步优势 actor-critic(Asynchronous Advantage Actor-Critic)算法,已成功用于在各种视频游戏中达到超人类表现。随着 Google Deepmind 与 Blizzard Entertainment 提出的 pysc2 学习环境的发布,星际争霸 II 成为强化学习界的新挑战。尽管是若干 AI 开发者的目标,但鲜有达到人类水平表现者。在本项目中,我们解释该环境的复杂性并讨论我们在该环境上实验的结果。我们比较了多种架构,并证明迁移学习可成为需要技能迁移的复杂场景下强化学习研究的有效范式。
引用
@article{arxiv.1807.08217,
title = {Asynchronous Advantage Actor-Critic Agent for Starcraft II},
author = {Basel Alghanem and Keerthana P G},
journal= {arXiv preprint arXiv:1807.08217},
year = {2018}
}
备注
arXiv admin note: text overlap with arXiv:1708.04782 by other authors