中文
相关论文

相关论文: Convergence Guarantees of Policy Optimization Meth…

200 篇论文

We suggest and compare different methods for the numerical solution of Lyapunov like equations with application to control of Markovian jump linear systems. First, we consider fixed point iterations and associated Krylov subspace…

数值分析 · 数学 2017-03-14 Tobias Damm , Kazuhiro Sato , Axel Vierling

We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-quadratic systems, and study various settings of driving…

机器学习 · 计算机科学 2020-05-19 Dhruv Malik , Ashwin Pananjady , Kush Bhatia , Koulik Khamaru , Peter L. Bartlett , Martin J. Wainwright

Learning how to effectively control unknown dynamical systems is crucial for intelligent autonomous systems. This task becomes a significant challenge when the underlying dynamics are changing with time. Motivated by this challenge, this…

机器学习 · 计算机科学 2025-10-21 Yahya Sattar , Zhe Du , Davoud Ataee Tarzanagh , Laura Balzano , Necmiye Ozay , Samet Oymak

Feedback controllers for port-Hamiltonian systems reveal an intrinsic inverse optimality property since each passivating state feedback controller is optimal with respect to some specific performance index. Due to the nonlinear…

最优化与控制 · 数学 2020-07-20 Lukas Kölsch , Pol Jané Soneira , Felix Strehle , Sören Hohmann

In this paper, we focus on formal synthesis of control policies for finite Markov decision processes with non-negative real-valued costs. We develop an algorithm to automatically generate a policy that guarantees the satisfaction of a…

计算机科学中的逻辑 · 计算机科学 2013-09-10 Maria Svorenova , Ivana Cerna , Calin Belta

We study the policy gradient method (PGM) for the linear quadratic Gaussian (LQG) dynamic output-feedback control problem using an input-output-history (IOH) representation of the closed-loop system. First, we show that any dynamic…

最优化与控制 · 数学 2025-10-23 Tomonori Sadamoto , Takashi Tanaka

This paper proposes a stochastic model predictive control method for linear systems affected by additive Gaussian disturbances that optimizes over disturbance feedback matrices online. Closed-loop satisfaction of probabilistic constraints…

系统与控制 · 电气工程与系统科学 2026-02-03 Marcell Bartos , Alexandre Didier , Jerome Sieber , Johannes Köhler , Melanie N. Zeilinger

We consider the continuous-time Linear-Quadratic-Regulator (LQR) problem in terms of optimizing a real-valued matrix function over the set of feedback gains. The results developed are in parallel to those in Bu et al. [1] for discrete-time…

系统与控制 · 电气工程与系统科学 2020-06-17 Jingjing Bu , Afshin Mesbahi , Mehran Mesbahi

Natural policy gradient (NPG) methods are among the most widely used policy optimization algorithms in contemporary reinforcement learning. This class of methods is often applied in conjunction with entropy regularization -- an algorithmic…

机器学习 · 统计学 2022-09-13 Shicong Cen , Chen Cheng , Yuxin Chen , Yuting Wei , Yuejie Chi

This paper addresses the problem of robust and optimal control for the class of nonlinear quadratic systems subject to norm-bounded parametric uncertainties and disturbances, and in presence of some amplitude constraints on the control…

系统与控制 · 计算机科学 2017-01-12 Merola Alessio , Cosentino Carlo , Colacino Domenico , Amato Francesco

In this paper, we consider stochastic optimal control of Markov Jump Linear Systems with state feedback but without observation of the jumping parameter. The proposed control law is assumed to be linear with constant gains that can be…

系统与控制 · 计算机科学 2015-07-02 Maxim Dolgov , Uwe D. Hanebeck

Linear dynamical systems that obey stochastic differential equations are canonical models. While optimal control of known systems has a rich literature, the problem is technically hard under model uncertainty and there are hardly any…

系统与控制 · 电气工程与系统科学 2023-06-09 Mohamad Kazem Shirani Faradonbeh , Mohamad Sadegh Shirani Faradonbeh

Incorporating pattern-learning for prediction (PLP) in many discrete-time or discrete-event systems allows for computation-efficient controller design by memorizing patterns to schedule control policies based on their future occurrences. In…

系统与控制 · 电气工程与系统科学 2023-05-10 SooJean Han , Soon-Jo Chung , John C. Doyle

We present a model-based globally convergent policy gradient method (PGM) for linear quadratic Gaussian (LQG) control. Firstly, we establish equivalence between optimizing dynamic output feedback controllers and designing a static feedback…

最优化与控制 · 数学 2024-02-27 Tomonori Sadamoto , Fumiya Nakamata

In this paper, we solve the chance-constrained covariance steering problem for discrete-time Markov Jump Linear Systems (MJLS) using a convex optimization framework. We derive the analytical expressions for the mean and covariance…

最优化与控制 · 数学 2026-05-28 Shaurya Shrivastava , Kenshiro Oguri

We consider discrete-time infinite horizon deterministic optimal control problems with nonnegative cost per stage, and a destination that is cost-free and absorbing. The classical linear-quadratic regulator problem is a special case. Our…

最优化与控制 · 数学 2017-12-20 Dimitri P. Bertsekas

We consider the problem of optimally controlling stochastic, Markovian systems subject to joint chance constraints over a finite-time horizon. For such problems, standard Dynamic Programming is inapplicable due to the time correlation of…

最优化与控制 · 数学 2024-11-22 Niklas Schmid , Marta Fochesato , Sarah H. Q. Li , Tobias Sutter , John Lygeros

This paper is concerned with the linear quadratic optimal control problem for networked system simultaneously with input delay and Markovian dropout. Different from the results in the literature, we consider the hold-input strategy, which…

最优化与控制 · 数学 2020-10-16 Hongdan Li , Xun Li , Huanshui Zhang

Many of the recent trajectory optimization algorithms alternate between linear approximation of the system dynamics around the mean trajectory and conservative policy update. One way of constraining the policy change is by bounding the…

机器学习 · 计算机科学 2018-07-03 Riad Akrour , Abbas Abdolmaleki , Hany Abdulsamad , Jan Peters , Gerhard Neumann

Approximate Newton methods are a standard optimization tool which aim to maintain the benefits of Newton's method, such as a fast rate of convergence, whilst alleviating its drawbacks, such as computationally expensive calculation or…

人工智能 · 计算机科学 2015-08-07 Thomas Furmston , Guy Lever