中文

协作多智能体强化学习中动作值网络分解的分析

多智能体系统 2024-12-20 v4

摘要

近年来,深度强化学习技术被应用于协作多智能体系统,并取得了巨大的经验性成功。然而,由于缺乏理论洞察,仍不清楚所采用的神经网络在学习什么,或者我们应怎样增强其学习能力以解决它们失败的问题。本工作中,我们在一系列一次性博弈上经验性地研究了各种网络架构的学习能力。尽管这些博弈简单,却捕捉了多智能体设定中出现的许多关键问题,例如指数级数量的联合动作或缺乏显式协调机制。我们的结果扩展了文献[4]中的结果,并量化了各种方法表示所需值函数的能力,帮助我们识别可能阻碍良好性能的原因,如值的稀疏性或过紧的协调要求。

关键词

引用

@article{arxiv.1902.07497,
  title  = {Analysing Factorizations of Action-Value Networks for Cooperative Multi-Agent Reinforcement Learning},
  author = {Jacopo Castellini and Frans A. Oliehoek and Rahul Savani and Shimon Whiteson},
  journal= {arXiv preprint arXiv:1902.07497},
  year   = {2024}
}

备注

This work as been accepted as an Extended Abstract in Proc. of the 18th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2019), N. Agmon, M. E. Taylor, E. Elkind, M. Veloso (eds.), May 2019, Montreal, Canada