中文

通过冲突感知梯度调整实现混合动力博弈中的公平合作

多智能体系统 2025-08-26 v1

摘要

混合动力环境下的多智能体强化学习 presents a fundamental challenge: agents must balance individual interests with collective goals, which are neither fully aligned nor strictly opposed. 为此,已提出礼物和内在动机等奖励重构方法。然而,这些方法主要关注通过管理 individual 与 collective returns 之间的权衡来促进合作,未能明确针对 agents 特定任务奖励的公平性进行考虑。本文提出一种自适应冲突感知梯度调整方法,以在保证 individual 奖励公平性的前提下促进合作。该方法在 two objectives 产生冲突时动态平衡来自 individual 与 collective 目标的 policy 梯度。通过明确解决此类冲突,该方法在保持 agents 之间公平性的同时提升 collective performance。我们提供理论结果,保证 both collective 与 individual 目标均实现单调非降低,且确保公平性。在顺序社交困境环境中的实验结果表明,该方法在 social welfare 方面优于基线方法,同时确保 agents 之间的公平性。

关键词

引用

@article{arxiv.2508.17696,
  title  = {Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient Adjustment},
  author = {Woojun Kim and Katia Sycara},
  journal= {arXiv preprint arXiv:2508.17696},
  year   = {2025}
}