English

Logarithmic regret for episodic continuous-time linear-quadratic reinforcement learning over a finite-time horizon

Optimization and Control 2022-06-22 v4 Machine Learning Machine Learning

Abstract

We study finite-time horizon continuous-time linear-quadratic reinforcement learning problems in an episodic setting, where both the state and control coefficients are unknown to the controller. We first propose a least-squares algorithm based on continuous-time observations and controls, and establish a logarithmic regret bound of order O((lnM)(lnlnM))O((\ln M)(\ln\ln M)), with MM being the number of learning episodes. The analysis consists of two parts: perturbation analysis, which exploits the regularity and robustness of the associated Riccati differential equation; and parameter estimation error, which relies on sub-exponential properties of continuous-time least-squares estimators. We further propose a practically implementable least-squares algorithm based on discrete-time observations and piecewise constant controls, which achieves similar logarithmic regret with an additional term depending explicitly on the time stepsizes used in the algorithm.

Keywords

Cite

@article{arxiv.2006.15316,
  title  = {Logarithmic regret for episodic continuous-time linear-quadratic reinforcement learning over a finite-time horizon},
  author = {Matteo Basei and Xin Guo and Anran Hu and Yufei Zhang},
  journal= {arXiv preprint arXiv:2006.15316},
  year   = {2022}
}

Comments

In this version, we added some comments and references