中文

多智能体强化学习中基于不可靠内在奖励的探索

人工智能 2019-06-06 v1

摘要

本文研究了利用内在奖励引导多智能体强化学习中探索的方法。我们讨论了将内在奖励应用于多个协作智能体时所面临的挑战,并展示了不可靠奖励如何阻碍分散式智能体学习最优策略。我们提出了一个新颖的框架——独立集中辅助 Q 学习(ICQL)来解决该问题,其中分散式智能体与集中式智能体共享控制权与经验回放缓冲区。仅集中式智能体获得内在奖励,但分散式智能体仍受益于改进了的探索,而免受不可靠激励的干扰。

关键词

引用

@article{arxiv.1906.02138,
  title  = {Exploration with Unreliable Intrinsic Reward in Multi-Agent Reinforcement Learning},
  author = {Wendelin Böhmer and Tabish Rashid and Shimon Whiteson},
  journal= {arXiv preprint arXiv:1906.02138},
  year   = {2019}
}

备注

Accepted to the 2nd Exploration in Reinforcement Learning Workshop at the International Conference on Machine Learning 2019