多智能体强化学习中基于不可靠内在奖励的探索
人工智能
2019-06-06 v1
摘要
本文研究了利用内在奖励引导多智能体强化学习中探索的方法。我们讨论了将内在奖励应用于多个协作智能体时所面临的挑战,并展示了不可靠奖励如何阻碍分散式智能体学习最优策略。我们提出了一个新颖的框架——独立集中辅助 Q 学习(ICQL)来解决该问题,其中分散式智能体与集中式智能体共享控制权与经验回放缓冲区。仅集中式智能体获得内在奖励,但分散式智能体仍受益于改进了的探索,而免受不可靠激励的干扰。
引用
@article{arxiv.1906.02138,
title = {Exploration with Unreliable Intrinsic Reward in Multi-Agent Reinforcement Learning},
author = {Wendelin Böhmer and Tabish Rashid and Shimon Whiteson},
journal= {arXiv preprint arXiv:1906.02138},
year = {2019}
}
备注
Accepted to the 2nd Exploration in Reinforcement Learning Workshop at the International Conference on Machine Learning 2019