English
Related papers

Related papers: Convergence Guarantees of Policy Optimization Meth…

200 papers

We propose a general framework for entropy-regularized average-reward reinforcement learning in Markov decision processes (MDPs). Our approach is based on extending the linear-programming formulation of policy optimization in MDPs to…

Machine Learning · Computer Science 2017-05-23 Gergely Neu , Anders Jonsson , Vicenç Gómez

We present a learning model predictive control (MPC) scheme for chance-constrained Markov jump systems with unknown switching probabilities. Using samples of the underlying Markov chain, ambiguity sets of transition probabilities are…

Optimization and Control · Mathematics 2023-01-06 Mathijs Schuurmans , Panagiotis Patrinos

We study optimal investment strategies that maximize expected utility from consumption and terminal wealth in a pure-jump asset price model with Markov-modulated (regime switching) jump-size distributions. We give sufficient conditions for…

Portfolio Management · Quantitative Finance 2014-06-13 Oscar Lopez , Rafael Serrano

Reinforcement learning (RL) is currently one of the most prominent methods for optimizing dynamical systems, with breakthrough results across various fields. The framework is based on the concept of a Markov decision process (MDP), leading…

Optimization and Control · Mathematics 2025-11-17 Rene Carmona , Mathieu Lauriere

Following Demidovich's concept and definition of convergent systems, we analyze the optimal nonlinear damping control, recently proposed [1] for the second-order systems. Targeting the problem of output regulation, correspondingly tracking…

Systems and Control · Electrical Eng. & Systems 2021-06-03 Michael Ruderman

We develop an exactly solvable framework of Markov decision process with a finite horizon, and continuous state and action spaces. We first review the exact solution of conventional linear quadratic regulation with a linear transition and a…

Machine Learning · Computer Science 2020-12-16 Yuan Yao , Xiaolin Sun

We revisit the finite time analysis of policy gradient methods in the one of the simplest settings: finite state and action MDPs with a policy class consisting of all stochastic policies and with exact gradient evaluations. There has been…

Machine Learning · Computer Science 2021-12-14 Jalaj Bhandari , Daniel Russo

In this work we study the convergence of gradient methods for nonconvex optimization problems -- specifically the effect of the problem formulation to the convergence behavior of the solution of a gradient flow. We show through a simple…

Optimization and Control · Mathematics 2025-10-03 Moh Kamalul Wafi , Arthur Castello B. de Oliveira , Eduardo D. Sontag

The linear quadratic regulator is the fundamental problem of optimal control. Its state feedback version was set and solved in the early 1960s. However the static output feedback problem has no explicit-form solution. It is suggested to…

Optimization and Control · Mathematics 2020-11-03 Ilyas Fatkhullin , Boris Polyak

Microgrids have more operational flexibilities as well as uncertainties than conventional power grids, especially when renewable energy resources are utilized. An energy storage based feedback controller can compensate undesired dynamics of…

Systems and Control · Electrical Eng. & Systems 2022-03-10 Tianwei Xia , Kai Sun , Wei Kang

This paper studies the adaptive optimal stationary control of continuous-time linear stochastic systems with both additive and multiplicative noises, using reinforcement learning techniques. Based on policy iteration, a novel off-policy…

Systems and Control · Electrical Eng. & Systems 2021-12-07 Bo Pang , Zhong-Ping Jiang

The goal of this paper is to present a method for simultaneous trajectory and local stabilizing policy optimization to generate local policies for trajectory-centric model-based reinforcement learning (MBRL). This is motivated by the fact…

This paper thoroughly investigates stochastic linear-quadratic optimal control problems with the Markovian regime switching system, where the coefficients of the state equation and the weighting matrices of the cost functional are random.…

Optimization and Control · Mathematics 2022-08-03 Jiaqiang Wen , Xun Li , Jie Xiong , Xin Zhang

The goal of this paper is to analyze distributional Markov Decision Processes as a class of control problems in which the objective is to learn policies that steer the distribution of a cumulative reward toward a prescribed target law,…

Optimization and Control · Mathematics 2026-02-09 Nicole Bäuerle , Athanasios Vasileiadis

Reinforcement learning with multinomial logistic (MNL) function approximation has become an important framework due to its flexibility and broad applicability. While existing studies have established regret guarantees under worst-case…

Machine Learning · Statistics 2026-05-28 Wonyoung Kim , Min-Hwan Oh , Garud Iyengar , Assaf Zeevi

This paper studies the approximation of optimal control policies by quantized (discretized) policies for a very general class of Markov decision processes (MDPs). The problem is motivated by applications in networked control systems,…

Optimization and Control · Mathematics 2015-05-14 Naci Saldi , Serdar Yüksel , Tamás Linder

This paper presents a general description of a parameter estimation inverse problem for systems governed by nonlinear differential equations. The inverse problem is presented using optimal control tools with state constraints, where the…

Numerical Analysis · Mathematics 2018-06-28 Mohamed Kamel Riahi , Issam Al Qattan

Policy gradient methods are a powerful family of reinforcement learning algorithms for continuous control that optimize a policy directly. However, standard first-order methods often converge slowly. Second-order methods can accelerate…

Systems and Control · Electrical Eng. & Systems 2025-11-05 Amirreza Valaei , Arash Bahari Kordabad , Sadegh Soudjani

This paper addresses the inverse optimal control problem of finding the state weighting function that leads to a quadratic value function when the cost on the input is fixed to be quadratic. The paper focuses on a class of infinite horizon…

Optimization and Control · Mathematics 2022-11-21 Luis Rodrigues

We study the pointwise stabilizability of a discrete-time, time-homogeneous, and stationary Markovian jump linear system. By using measure theory, ergodic theory and a splitting theorem of state space we show in a relatively simple way that…

Probability · Mathematics 2013-09-02 Xiongping Dai , Yu Huang , Mingqing Xiao