English

Reinforcement Learning in Deep Structured Teams: Initial Results with Finite and Infinite Valued Features

Multiagent Systems 2021-02-09 v4

Abstract

In this paper, we consider Markov chain and linear quadratic models for deep structured teams with discounted and time-average cost functions under two non-classical information structures, namely, deep state sharing and no sharing. In deep structured teams, agents are coupled in dynamics and cost functions through deep state, where deep state refers to a set of orthogonal linear regressions of the states. In this article, we consider a homogeneous linear regression for Markov chain models (i.e., empirical distribution of states) and a few orthonormal linear regressions for linear quadratic models (i.e., weighted average of states). Some planning algorithms are developed for the case when the model is known, and some reinforcement learning algorithms are proposed for the case when the model is not known completely. The convergence of two model-free (reinforcement learning) algorithms, one for Markov chain models and one for linear quadratic models, is established. The results are then applied to a smart grid.

Keywords

Cite

@article{arxiv.2010.02868,
  title  = {Reinforcement Learning in Deep Structured Teams: Initial Results with Finite and Infinite Valued Features},
  author = {Jalal Arabneydi and Masoud Roudneshin and Amir G. Aghdam},
  journal= {arXiv preprint arXiv:2010.02868},
  year   = {2021}
}

Comments

This version corrects some typographical errors

R2 v1 2026-06-23T19:05:45.629Z