English
Related papers

Related papers: Convergence Guarantees of Policy Optimization Meth…

200 papers

Adaptive optimal control of nonlinear dynamic systems with deterministic and known dynamics under a known undiscounted infinite-horizon cost function is investigated. Policy iteration scheme initiated using a stabilizing initial control is…

Systems and Control · Computer Science 2015-05-21 Ali Heydari

Policy gradient methods, where one searches for the policy of interest by maximizing the value functions using first-order information, become increasingly popular for sequential decision making in reinforcement learning, games, and…

Optimization and Control · Mathematics 2023-10-10 Shicong Cen , Yuejie Chi

Policy gradient algorithms are widely used in reinforcement learning and belong to the class of approximate dynamic programming methods. This paper studies two key policy gradient algorithms, the Natural Policy Gradient and the Gauss-Newton…

Systems and Control · Electrical Eng. & Systems 2026-05-11 Bowen Song , Sebastien Gros , Andrea Iannelli

The article poses a general model for optimal control subject to information constraints, motivated in part by recent work of Sims and others on information-constrained decision-making by economic agents. In the average-cost optimal control…

Optimization and Control · Mathematics 2016-02-24 Ehsan Shafieepoorfard , Maxim Raginsky , Sean P. Meyn

A method for deriving accurate analytic approximations for Markovian open quantum systems was recently introduced in [F. Lucas and K. Hornberger, Phys. Rev. Lett. 110, 240401 (2013)]. Here, we present a detailed derivation of the underlying…

Quantum Physics · Physics 2014-03-10 Felix Lucas , Klaus Hornberger

We consider an improper reinforcement learning setting where a learner is given $M$ base controllers for an unknown Markov decision process, and wishes to combine them optimally to produce a potentially new controller that can outperform…

Machine Learning · Computer Science 2021-07-06 Mohammadi Zaki , Avinash Mohan , Aditya Gopalan , Shie Mannor

Direct policy search has achieved great empirical success in reinforcement learning. Recently, there has been increasing interest in studying its theoretical properties for continuous control, and fruitful results have been established for…

Optimization and Control · Mathematics 2023-04-04 Yujie Tang , Yang Zheng

We propose a globally convergent Gauss-Newton algorithm for finding a local optimal solution of a non-convex and possibly non-smooth optimization problem. The algorithm that we present is based on a Gauss-Newton-type iteration for the…

Optimization and Control · Mathematics 2020-12-08 Ilyes Mezghani , Quoc Tran-Dinh , Ion Necoara , Anthony Papavasiliou

Iterative trajectory optimization techniques for non-linear dynamical systems are among the most powerful and sample-efficient methods of model-based reinforcement learning and approximate optimal control. By leveraging time-variant local…

Systems and Control · Electrical Eng. & Systems 2019-08-01 Onur Celik , Hany Abdulsamad , Jan Peters

Policy optimization methods with function approximation are widely used in multi-agent reinforcement learning. However, it remains elusive how to design such algorithms with statistical guarantees. Leveraging a multi-agent performance…

Machine Learning · Computer Science 2023-05-09 Yulai Zhao , Zhuoran Yang , Zhaoran Wang , Jason D. Lee

This paper presents a theoretical overview of a Neural Contraction Metric (NCM): a neural network model of an optimal contraction metric and corresponding differential Lyapunov function, the existence of which is a necessary and sufficient…

Machine Learning · Computer Science 2021-10-05 Hiroyasu Tsukamoto , Soon-Jo Chung , Jean-Jacques Slotine , Chuchu Fan

We propose a method for approximating solutions to optimization problems involving the global stability properties of parameter-dependent continuous-time autonomous dynamical systems. The method relies on an approximation of the…

Optimization and Control · Mathematics 2013-08-12 Péter Koltai , Alexander Volf

We study the convergence of an $N$-particle Markovian controlled system to the solution of a family of stochastic McKean-Vlasov control problems, either with a finite horizon or Schr\"odinger type cost functional. Specifically, under…

Probability · Mathematics 2024-05-22 Francesco C. De Vecchi , Chiara Rigoni

Model predictive control (MPC) is widely used in process control due to its interpretability and ability to handle constraints. As a parametric policy in reinforcement learning (RL), MPC offers strong initial performance and low data…

Systems and Control · Electrical Eng. & Systems 2026-04-03 Dean Brandner , Sebastien Gros , Sergio Lucia

In this paper, we present an equivalent convex optimization formulation for discrete-time stochastic linear systems subject to linear chance constraints, alongside a tight convex relaxation for quadratic chance constraints. By lifting the…

Systems and Control · Electrical Eng. & Systems 2026-03-23 Tanmay Dokania , Yashwanth Kumar Nakka

Multi-Agent Reinforcement Learning (MARL) -- where multiple agents learn to interact in a shared dynamic environment -- permeates across a wide range of critical applications. While there has been substantial progress on understanding the…

Computer Science and Game Theory · Computer Science 2022-10-05 Shicong Cen , Yuejie Chi , Simon S. Du , Lin Xiao

In recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-feedback control,…

Optimization and Control · Mathematics 2023-10-31 Jingliang Duan , Jie Li , Xuyang Chen , Kai Zhao , Shengbo Eben Li , Lin Zhao

In most real cases transition probabilities between operational modes of Markov jump linear systems cannot be computed exactly and are time-varying. We take into account this aspect by considering Markov jump linear systems where the…

Systems and Control · Computer Science 2021-03-22 Y. Zacchia Lun , A. Abate , A. D'Innocenzo

Markov decision problems are most commonly solved via dynamic programming. Another approach is Bellman residual minimization, which directly minimizes the squared Bellman residual objective function. However, compared to dynamic…

Machine Learning · Computer Science 2026-04-28 Donghwan Lee , Hyukjun Yang

Automatic optimization of robotic behavior has been the long-standing goal of Evolutionary Robotics. Allowing the problem at hand to be solved by automation often leads to novel approaches and new insights. A common problem encountered with…

Robotics · Computer Science 2019-12-18 Kirk Y. W. Scheper , Guido C. H. E. de Croon
‹ Prev 1 4 5 6 7 8 10 Next ›