中文
相关论文

相关论文: Convergence Guarantees of Policy Optimization Meth…

200 篇论文

Adaptive optimal control of nonlinear dynamic systems with deterministic and known dynamics under a known undiscounted infinite-horizon cost function is investigated. Policy iteration scheme initiated using a stabilizing initial control is…

系统与控制 · 计算机科学 2015-05-21 Ali Heydari

Policy gradient methods, where one searches for the policy of interest by maximizing the value functions using first-order information, become increasingly popular for sequential decision making in reinforcement learning, games, and…

最优化与控制 · 数学 2023-10-10 Shicong Cen , Yuejie Chi

Policy gradient algorithms are widely used in reinforcement learning and belong to the class of approximate dynamic programming methods. This paper studies two key policy gradient algorithms, the Natural Policy Gradient and the Gauss-Newton…

系统与控制 · 电气工程与系统科学 2026-05-11 Bowen Song , Sebastien Gros , Andrea Iannelli

The article poses a general model for optimal control subject to information constraints, motivated in part by recent work of Sims and others on information-constrained decision-making by economic agents. In the average-cost optimal control…

最优化与控制 · 数学 2016-02-24 Ehsan Shafieepoorfard , Maxim Raginsky , Sean P. Meyn

A method for deriving accurate analytic approximations for Markovian open quantum systems was recently introduced in [F. Lucas and K. Hornberger, Phys. Rev. Lett. 110, 240401 (2013)]. Here, we present a detailed derivation of the underlying…

量子物理 · 物理学 2014-03-10 Felix Lucas , Klaus Hornberger

We consider an improper reinforcement learning setting where a learner is given $M$ base controllers for an unknown Markov decision process, and wishes to combine them optimally to produce a potentially new controller that can outperform…

机器学习 · 计算机科学 2021-07-06 Mohammadi Zaki , Avinash Mohan , Aditya Gopalan , Shie Mannor

Direct policy search has achieved great empirical success in reinforcement learning. Recently, there has been increasing interest in studying its theoretical properties for continuous control, and fruitful results have been established for…

最优化与控制 · 数学 2023-04-04 Yujie Tang , Yang Zheng

We propose a globally convergent Gauss-Newton algorithm for finding a local optimal solution of a non-convex and possibly non-smooth optimization problem. The algorithm that we present is based on a Gauss-Newton-type iteration for the…

最优化与控制 · 数学 2020-12-08 Ilyes Mezghani , Quoc Tran-Dinh , Ion Necoara , Anthony Papavasiliou

Iterative trajectory optimization techniques for non-linear dynamical systems are among the most powerful and sample-efficient methods of model-based reinforcement learning and approximate optimal control. By leveraging time-variant local…

系统与控制 · 电气工程与系统科学 2019-08-01 Onur Celik , Hany Abdulsamad , Jan Peters

Policy optimization methods with function approximation are widely used in multi-agent reinforcement learning. However, it remains elusive how to design such algorithms with statistical guarantees. Leveraging a multi-agent performance…

机器学习 · 计算机科学 2023-05-09 Yulai Zhao , Zhuoran Yang , Zhaoran Wang , Jason D. Lee

This paper presents a theoretical overview of a Neural Contraction Metric (NCM): a neural network model of an optimal contraction metric and corresponding differential Lyapunov function, the existence of which is a necessary and sufficient…

机器学习 · 计算机科学 2021-10-05 Hiroyasu Tsukamoto , Soon-Jo Chung , Jean-Jacques Slotine , Chuchu Fan

We propose a method for approximating solutions to optimization problems involving the global stability properties of parameter-dependent continuous-time autonomous dynamical systems. The method relies on an approximation of the…

最优化与控制 · 数学 2013-08-12 Péter Koltai , Alexander Volf

We study the convergence of an $N$-particle Markovian controlled system to the solution of a family of stochastic McKean-Vlasov control problems, either with a finite horizon or Schr\"odinger type cost functional. Specifically, under…

概率论 · 数学 2024-05-22 Francesco C. De Vecchi , Chiara Rigoni

Model predictive control (MPC) is widely used in process control due to its interpretability and ability to handle constraints. As a parametric policy in reinforcement learning (RL), MPC offers strong initial performance and low data…

系统与控制 · 电气工程与系统科学 2026-04-03 Dean Brandner , Sebastien Gros , Sergio Lucia

In this paper, we present an equivalent convex optimization formulation for discrete-time stochastic linear systems subject to linear chance constraints, alongside a tight convex relaxation for quadratic chance constraints. By lifting the…

系统与控制 · 电气工程与系统科学 2026-03-23 Tanmay Dokania , Yashwanth Kumar Nakka

Multi-Agent Reinforcement Learning (MARL) -- where multiple agents learn to interact in a shared dynamic environment -- permeates across a wide range of critical applications. While there has been substantial progress on understanding the…

计算机科学与博弈论 · 计算机科学 2022-10-05 Shicong Cen , Yuejie Chi , Simon S. Du , Lin Xiao

In recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-feedback control,…

最优化与控制 · 数学 2023-10-31 Jingliang Duan , Jie Li , Xuyang Chen , Kai Zhao , Shengbo Eben Li , Lin Zhao

In most real cases transition probabilities between operational modes of Markov jump linear systems cannot be computed exactly and are time-varying. We take into account this aspect by considering Markov jump linear systems where the…

系统与控制 · 计算机科学 2021-03-22 Y. Zacchia Lun , A. Abate , A. D'Innocenzo

Markov decision problems are most commonly solved via dynamic programming. Another approach is Bellman residual minimization, which directly minimizes the squared Bellman residual objective function. However, compared to dynamic…

机器学习 · 计算机科学 2026-04-28 Donghwan Lee , Hyukjun Yang

Automatic optimization of robotic behavior has been the long-standing goal of Evolutionary Robotics. Allowing the problem at hand to be solved by automation often leads to novel approaches and new insights. A common problem encountered with…

机器人学 · 计算机科学 2019-12-18 Kirk Y. W. Scheper , Guido C. H. E. de Croon