English

Meta-Learning Linear Quadratic Regulators: A Policy Gradient MAML Approach for Model-free LQR

Optimization and Control 2024-06-04 v2 Machine Learning

Abstract

We investigate the problem of learning linear quadratic regulators (LQR) in a multi-task, heterogeneous, and model-free setting. We characterize the stability and personalization guarantees of a policy gradient-based (PG) model-agnostic meta-learning (MAML) (Finn et al., 2017) approach for the LQR problem under different task-heterogeneity settings. We show that our MAML-LQR algorithm produces a stabilizing controller close to each task-specific optimal controller up to a task-heterogeneity bias in both model-based and model-free learning scenarios. Moreover, in the model-based setting, we show that such a controller is achieved with a linear convergence rate, which improves upon sub-linear rates from existing work. Our theoretical guarantees demonstrate that the learned controller can efficiently adapt to unseen LQR tasks.

Keywords

Cite

@article{arxiv.2401.14534,
  title  = {Meta-Learning Linear Quadratic Regulators: A Policy Gradient MAML Approach for Model-free LQR},
  author = {Leonardo F. Toso and Donglin Zhan and James Anderson and Han Wang},
  journal= {arXiv preprint arXiv:2401.14534},
  year   = {2024}
}
R2 v1 2026-06-28T14:27:37.526Z