English

Learning Decentralized Linear Quadratic Regulators with $\sqrt{T}$ Regret

Optimization and Control 2024-07-08 v4 Machine Learning Systems and Control Systems and Control

Abstract

We propose an online learning algorithm that adaptively designs a decentralized linear quadratic regulator when the system model is unknown a priori and new data samples from a single system trajectory become progressively available. The algorithm uses a disturbance-feedback representation of state-feedback controllers coupled with online convex optimization with memory and delayed feedback. Under the assumption that the system is stable or given a known stabilizing controller, we show that our controller enjoys an expected regret that scales as T\sqrt{T} with the time horizon TT for the case of partially nested information pattern. For more general information patterns, the optimal controller is unknown even if the system model is known. In this case, the regret of our controller is shown with respect to a linear sub-optimal controller. We validate our theoretical findings using numerical experiments.

Keywords

Cite

@article{arxiv.2210.08886,
  title  = {Learning Decentralized Linear Quadratic Regulators with $\sqrt{T}$ Regret},
  author = {Lintao Ye and Ming Chi and Ruiquan Liao and Vijay Gupta},
  journal= {arXiv preprint arXiv:2210.08886},
  year   = {2024}
}

Comments

50 pages, 3 figures

R2 v1 2026-06-28T03:47:39.558Z