English

Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning

Multiagent Systems 2024-12-20 v1 Artificial Intelligence Computer Science and Game Theory Machine Learning Robotics

Abstract

In multi-agent environments, agents often struggle to learn optimal policies due to sparse or delayed global rewards, particularly in long-horizon tasks where it is challenging to evaluate actions at intermediate time steps. We introduce Temporal-Agent Reward Redistribution (TAR2^2), a novel approach designed to address the agent-temporal credit assignment problem by redistributing sparse rewards both temporally and across agents. TAR2^2 decomposes sparse global rewards into time-step-specific rewards and calculates agent-specific contributions to these rewards. We theoretically prove that TAR2^2 is equivalent to potential-based reward shaping, ensuring that the optimal policy remains unchanged. Empirical results demonstrate that TAR2^2 stabilizes and accelerates the learning process. Additionally, we show that when TAR2^2 is integrated with single-agent reinforcement learning algorithms, it performs as well as or better than traditional multi-agent reinforcement learning methods.

Keywords

Cite

@article{arxiv.2412.14779,
  title  = {Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning},
  author = {Aditya Kapoor and Sushant Swamy and Kale-ab Tessera and Mayank Baranwal and Mingfei Sun and Harshad Khadilkar and Stefano V. Albrecht},
  journal= {arXiv preprint arXiv:2412.14779},
  year   = {2024}
}

Comments

12 pages, 1 figure

R2 v1 2026-06-28T20:42:07.488Z