English

Non-Stationary Policy Learning for Multi-Timescale Multi-Agent Reinforcement Learning

Machine Learning 2023-07-19 v1 Artificial Intelligence Multiagent Systems Systems and Control Systems and Control

Abstract

In multi-timescale multi-agent reinforcement learning (MARL), agents interact across different timescales. In general, policies for time-dependent behaviors, such as those induced by multiple timescales, are non-stationary. Learning non-stationary policies is challenging and typically requires sophisticated or inefficient algorithms. Motivated by the prevalence of this control problem in real-world complex systems, we introduce a simple framework for learning non-stationary policies for multi-timescale MARL. Our approach uses available information about agent timescales to define a periodic time encoding. In detail, we theoretically demonstrate that the effects of non-stationarity introduced by multiple timescales can be learned by a periodic multi-agent policy. To learn such policies, we propose a policy gradient algorithm that parameterizes the actor and critic with phase-functioned neural networks, which provide an inductive bias for periodicity. The framework's ability to effectively learn multi-timescale policies is validated on a gridworld and building energy management environment.

Keywords

Cite

@article{arxiv.2307.08794,
  title  = {Non-Stationary Policy Learning for Multi-Timescale Multi-Agent Reinforcement Learning},
  author = {Patrick Emami and Xiangyu Zhang and David Biagioni and Ahmed S. Zamzam},
  journal= {arXiv preprint arXiv:2307.08794},
  year   = {2023}
}

Comments

Accepted at IEEE CDC'23. 7 pages, 6 figures

R2 v1 2026-06-28T11:32:55.973Z