English

On Information Asymmetry in Competitive Multi-Agent Reinforcement Learning: Convergence and Optimality

Machine Learning 2021-01-26 v2 Computer Science and Game Theory Multiagent Systems Systems and Control Theoretical Economics Systems and Control

Abstract

In this work, we study the system of interacting non-cooperative two Q-learning agents, where one agent has the privilege of observing the other's actions. We show that this information asymmetry can lead to a stable outcome of population learning, which generally does not occur in an environment of general independent learners. The resulting post-learning policies are almost optimal in the underlying game sense, i.e., they form a Nash equilibrium. Furthermore, we propose in this work a Q-learning algorithm, requiring predictive observation of two subsequent opponent's actions, yielding an optimal strategy given that the latter applies a stationary strategy, and discuss the existence of the Nash equilibrium in the underlying information asymmetrical game.

Keywords

Cite

@article{arxiv.2010.10901,
  title  = {On Information Asymmetry in Competitive Multi-Agent Reinforcement Learning: Convergence and Optimality},
  author = {Ezra Tampubolon and Haris Ceribasic and Holger Boche},
  journal= {arXiv preprint arXiv:2010.10901},
  year   = {2021}
}

Comments

Preprint

R2 v1 2026-06-23T19:31:05.774Z