English
Related papers

Related papers: Convergence of Policy Gradient for Stochastic Line…

200 papers

Recently, policy optimization for control purposes has received renewed attention due to the increasing interest in reinforcement learning. In this paper, we investigate the global convergence of gradient-based policy optimization methods…

Optimization and Control · Mathematics 2020-11-25 Joao Paulo Jansch-Porto , Bin Hu , Geir Dullerud

We study a signature-driven numerical scheme to solve multi-dimensional linear-quadratic (LQ) stochastic control problems. Using that linear signature functionals are dense in the natural class of admissible controls, we show that our…

Optimization and Control · Mathematics 2026-03-02 Alif Aqsha , Peter Bank , Leandro Sánchez-Betancourt

Off-policy evaluation of sequential decision policies from observational data is necessary in applications of batch reinforcement learning such as education and healthcare. In such settings, however, unobserved variables confound observed…

Machine Learning · Computer Science 2020-07-14 Nathan Kallus , Angela Zhou

We prove that, for finite-arm bandits with linear function approximation, the global convergence of policy gradient (PG) methods depends on inter-related properties between the policy update and the representation. textcolor{blue}{First},…

Machine Learning · Computer Science 2025-04-04 Jincheng Mei , Bo Dai , Alekh Agarwal , Mohammad Ghavamzadeh , Csaba Szepesvari , Dale Schuurmans

We propose a gradient-based method for quadratic programming problems with a single linear constraint and bounds on the variables. Inspired by the GPCG algorithm for bound-constrained convex quadratic programming [J.J. Mor\'e and G.…

Optimization and Control · Mathematics 2019-02-19 Daniela di Serafino , Gerardo Toraldo , Marco Viola , Jesse Barlow

We study a new two-time-scale stochastic gradient method for solving optimization problems, where the gradients are computed with the aid of an auxiliary variable under samples generated by time-varying MDPs controlled by the underlying…

Optimization and Control · Mathematics 2024-08-27 Sihan Zeng , Thinh T. Doan , Justin Romberg

Proximal gradient methods are popular in sparse optimization as they are straightforward to implement. Nevertheless, they achieve biased solutions, requiring many iterations to converge. This work addresses these issues through a suitable…

Optimization and Control · Mathematics 2025-04-18 V. Cerone , S. M. Fosson , A. Re , D. Regruto

Optimal control problems with a very large time horizon can be tackled with the Receding Horizon Control (RHC) method, which consists in solving a sequence of optimal control problems with small prediction horizon. The main result of this…

Optimization and Control · Mathematics 2020-02-03 Tobias Breiten , Laurent Pfeiffer

We propose two policy gradient algorithms for solving the problem of control in an off-policy reinforcement learning (RL) context. Both algorithms incorporate a smoothed functional (SF) based gradient estimation scheme. The first algorithm…

Machine Learning · Computer Science 2024-06-25 Nithia Vijayan , Prashanth L. A

We study the problem of adaptive control of the stochastic linear quadratic regulator (LQR) with constraints that must be satisfied at every time step. Prior work on the multidimensional problem has shown $\tilde{O}(T^{2/3})$ regret and…

Optimization and Control · Mathematics 2026-05-08 Spencer Hutchinson , Nanfei Jiang , Mahnoosh Alizadeh

We analyze the convergence rate of the unregularized natural policy gradient algorithm with log-linear policy parametrizations in infinite-horizon discounted Markov decision processes. In the deterministic case, when the Q-value is known…

Machine Learning · Computer Science 2023-03-15 Carlo Alfano , Patrick Rebeschini

We consider the problem of computing optimal linear control policies for linear systems in finite-horizon. The states and the inputs are required to remain inside pre-specified safety sets at all times despite unknown disturbances. In this…

Systems and Control · Computer Science 2019-12-17 Luca Furieri , Maryam Kamgarpour

This work presents an algorithmic scheme for solving the infinite-time constrained linear quadratic regulation problem. We employ an accelerated version of a popular proximal gradient scheme, commonly known as the Forward-Backward Splitting…

Optimization and Control · Mathematics 2015-01-20 Giorgos Stathopoulos , Milan Korda , Colin N. Jones

Stabilizing a dynamical system is a fundamental problem that serves as a cornerstone for many complex tasks in the field of control systems. The problem becomes challenging when the system model is unknown. Among the Reinforcement Learning…

Systems and Control · Electrical Eng. & Systems 2026-01-30 Ankang Zhang , Ming Chi , Xiaoling Wang , Lintao Ye

In order to model risk aversion in reinforcement learning, an emerging line of research adapts familiar algorithms to optimize coherent risk functionals, a class that includes conditional value-at-risk (CVaR). Because optimizing the…

Machine Learning · Computer Science 2021-03-09 Audrey Huang , Liu Leqi , Zachary C. Lipton , Kamyar Azizzadenesheli

Stochastic gradient descent is one of the most successful approaches for solving large-scale problems, especially in machine learning and statistics. At each iteration, it employs an unbiased estimator of the full gradient computed from one…

Numerical Analysis · Mathematics 2018-12-05 Bangti Jin , Xiliang Lu

This paper is concerned with a discounted optimal control problem of partially observed forward-backward stochastic systems with jumps on infinite horizon. The control domain is convex and a kind of infinite horizon observation equation is…

Optimization and Control · Mathematics 2022-01-04 Yueyang Zheng , Jingtao Shi

We study an optimal control problem on infinite time horizon with semimartingale strategies, random coefficients and regime switching. The value function and the optimal strategy can be characterized in terms of three systems of backward…

Optimization and Control · Mathematics 2026-02-27 Xinman Cheng , Guanxing Fu , Xiaonyu Xia

We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework…

Optimization and Control · Mathematics 2025-06-11 Qi Feng , Gu Wang

This paper is concerned with a discounted stochastic optimal control problem for regime switching diffusion in an infinite horizon. First, as a preliminary with particular interests in its own right, the global well-posedness of infinite…

Optimization and Control · Mathematics 2026-02-06 Kai Ding , Xun Li , Siyu Lv , Xin Zhang