中文

面向 Networked 代理的分布式价值分解网络

机器学习 2025-02-12 v1 人工智能 多智能体系统

摘要

我们研究了在部分可观测性下的分布式训练问题,即合作多智能体强化学习 (MARL) 代理 maximizes 预期累积联合奖励。我们提出了分布式价值分解网络 (DVDN),其生成分解为代理级 Q 函数的联合 Q 函数。与原始价值分解网络依赖集中训练不同,我们的方法适用于集中训练不可行且代理必须通过与物理环境的 Decentralized 交互学习的情境,同时与其同行通信。DVDN 通过局部估计共享目标来克服集中训练的需求。我们针对异构和同构代理设置分别贡献了两个创新算法,即 DVDN 和 DVDN (GT)。实验结果表明,尽管通信过程中存在信息损失,但两种算法在三个标准环境中的十个 MARL 任务中都能近似价值分解网络的性能。

关键词

引用

@article{arxiv.2502.07635,
  title  = {Distributed Value Decomposition Networks with Networked Agents},
  author = {Guilherme S. Varela and Alberto Sardinha and Francisco S. Melo},
  journal= {arXiv preprint arXiv:2502.07635},
  year   = {2025}
}

备注

21 pages, 15 figures, to be published in Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2025), Detroit, Michigan, USA, May 19 - 23, 2025, IFAAMAS