English
Related papers

Related papers: Convergence and Sample Complexity of Policy Gradie…

200 papers

We address the problem of designing stabilizing control policies for nonlinear systems in discrete-time, while minimizing an arbitrary cost function. When the system is linear and the cost is convex, the System Level Synthesis (SLS)…

Systems and Control · Electrical Eng. & Systems 2023-01-03 Luca Furieri , Clara Lucía Galimberti , Giancarlo Ferrari-Trecate

In this work we study the convergence of gradient methods for nonconvex optimization problems -- specifically the effect of the problem formulation to the convergence behavior of the solution of a gradient flow. We show through a simple…

Optimization and Control · Mathematics 2025-10-03 Moh Kamalul Wafi , Arthur Castello B. de Oliveira , Eduardo D. Sontag

Standard model-based control design deteriorates when the system dynamics change during operation. To overcome this challenge, online and adaptive methods have been proposed in the literature. In this work, we consider the class of…

Systems and Control · Electrical Eng. & Systems 2026-04-16 Marcell Bartos , Johannes Köhler , Florian Dörfler , Melanie N. Zeilinger

This paper addresses the end-to-end sample complexity bound for learning the H2 optimal controller (the Linear Quadratic Gaussian (LQG) problem) with unknown dynamics, for potentially unstable Linear Time Invariant (LTI) systems. The robust…

Systems and Control · Electrical Eng. & Systems 2022-07-05 Yifei Zhang , Sourav Kumar Ukil , Ephraim Neimand , Serban Sabau , Myron E. Hohil

Nonlinear control systems with partial information to the decision maker are prevalent in a variety of applications. As a step toward studying such nonlinear systems, this work explores reinforcement learning methods for finding the optimal…

Machine Learning · Computer Science 2025-04-11 Yinbin Han , Meisam Razaviyayn , Renyuan Xu

We present a model-based globally convergent policy gradient method (PGM) for linear quadratic Gaussian (LQG) control. Firstly, we establish equivalence between optimizing dynamic output feedback controllers and designing a static feedback…

Optimization and Control · Mathematics 2024-02-27 Tomonori Sadamoto , Fumiya Nakamata

In safety-critical applications, reinforcement learning (RL) needs to consider safety constraints. However, theoretical understandings of constrained RL for continuous control are largely absent. As a case study, this paper presents a…

Optimization and Control · Mathematics 2024-06-07 Feiran Zhao , Keyou You

The Linear Quadratic Gaussian (LQG) controller is known to be inherently fragile to model misspecifications common in real-world situations. We consider discrete-time partially observable stochastic linear systems and provide a…

Optimization and Control · Mathematics 2025-07-31 Marta Fochesato , Lucia Falconi , Mattia Zorzi , Augusto Ferrante , John Lygeros

We consider the optimal regulation problem for nonlinear control-affine dynamical systems. Whereas the linear-quadratic regulator (LQR) considers optimal control of a linear system with quadratic cost function, we study polynomial systems…

Optimization and Control · Mathematics 2024-10-30 Nicholas A. Corbin , Boris Kramer

Direct policy gradient methods for reinforcement learning and continuous control problems are a popular approach for a variety of reasons: 1) they are easy to implement without explicit knowledge of the underlying model 2) they are an…

Machine Learning · Computer Science 2019-03-26 Maryam Fazel , Rong Ge , Sham M. Kakade , Mehran Mesbahi

This paper introduces an extension of the LQR-tree algorithm, which is a feedback-motion-planning algorithm for stabilizing a system of ordinary differential equations from a bounded set of initial conditions to a goal. The constructed…

Systems and Control · Electrical Eng. & Systems 2023-03-02 Jiří Fejlek , Stefan Ratschan

We investigate the problem of learning an $\epsilon$-approximate solution for the discrete-time Linear Quadratic Regulator (LQR) problem via a Stochastic Variance-Reduced Policy Gradient (SVRPG) approach. Whilst policy gradient methods have…

Optimization and Control · Mathematics 2023-09-20 Leonardo F. Toso , Han Wang , James Anderson

Policy optimization (PO) is a key ingredient for reinforcement learning (RL). For control design, certain constraints are usually enforced on the policies to optimize, accounting for either the stability, robustness, or safety concerns on…

Optimization and Control · Mathematics 2021-02-16 Kaiqing Zhang , Bin Hu , Tamer Başar

We present a data-driven method for solving the linear quadratic regulator problem for systems with multiplicative disturbances, the distribution of which is only known through sample estimates. We adopt a distributionally robust approach…

Systems and Control · Electrical Eng. & Systems 2020-05-27 Peter Coppens , Mathijs Schuurmans , Panagiotis Patrinos

We propose and analyze a stabilizing iteration scheme for the algorithmic implementation of model predictive control for linear discrete-time systems. Polytopic input and state constraints are considered and handled by means of so-called…

Optimization and Control · Mathematics 2016-04-07 Christian Feller , Christian Ebenbauer

We consider the policy gradient adaptive control (PGAC) framework, which adaptively updates a control policy in real time, by performing data-based gradient descent steps on the linear quadratic regulator cost. This method has empirically…

Optimization and Control · Mathematics 2026-01-07 Felix Laurent , Feiran Zhao , Jaap Eising , Florian Dörfler

This paper presents a robust fixed lag smoother for a class of nonlinear uncertain systems. A unified scheme, which combines a nonlinear robust estimator with a stable fixed lag smoother, is presented to improve the error covariance of the…

Systems and Control · Computer Science 2013-09-10 Obaid Ur Rehman , Ian R. Petersen

Unlike traditional model-based reinforcement learning approaches that estimate system parameters from data, non-model-based data-driven control learns the optimal policy directly from input-state data without any intermediate model…

Optimization and Control · Mathematics 2026-05-05 Leilei Cui , Zhong-Ping Jiang , Petter N. Kolm , Grégoire G. Macqueron

This paper studies the robustness of policy iteration in the context of continuous-time infinite-horizon linear quadratic regulation (LQR) problem. It is shown that Kleinman's policy iteration algorithm is inherently robust to small…

Systems and Control · Electrical Eng. & Systems 2020-09-01 Bo Pang , Tao Bian , Zhong-Ping Jiang

Entropy regularization is an efficient technique for encouraging exploration and preventing a premature convergence of (vanilla) policy gradient methods in reinforcement learning (RL). However, the theoretical understanding of…

Machine Learning · Computer Science 2024-07-16 Yuhao Ding , Junzi Zhang , Hyunin Lee , Javad Lavaei