English

A relaxed technical assumption for posterior sampling-based reinforcement learning for control of unknown linear systems

Systems and Control 2022-09-21 v2 Artificial Intelligence Systems and Control Optimization and Control

Abstract

We revisit the Thompson sampling algorithm to control an unknown linear quadratic (LQ) system recently proposed by Ouyang et al (arXiv:1709.04047). The regret bound of the algorithm was derived under a technical assumption on the induced norm of the closed loop system. In this technical note, we show that by making a minor modification in the algorithm (in particular, ensuring that an episode does not end too soon), this technical assumption on the induced norm can be replaced by a milder assumption in terms of the spectral radius of the closed loop system. The modified algorithm has the same Bayesian regret of O~(T)\tilde{\mathcal{O}}(\sqrt{T}), where TT is the time-horizon and the O~()\tilde{\mathcal{O}}(\cdot) notation hides logarithmic terms in~TT.

Keywords

Cite

@article{arxiv.2108.08502,
  title  = {A relaxed technical assumption for posterior sampling-based reinforcement learning for control of unknown linear systems},
  author = {Mukul Gagrani and Sagar Sudhakara and Aditya Mahajan and Ashutosh Nayyar and Yi Ouyang},
  journal= {arXiv preprint arXiv:2108.08502},
  year   = {2022}
}
R2 v1 2026-06-24T05:14:31.798Z