English
Related papers

Related papers: Convergence and Sample Complexity of Policy Gradie…

200 papers

We introduce a generic solver for dynamic portfolio allocation problems when the market exhibits return predictability, price impact and partial observability. We assume that the price modeling can be encoded into a linear state-space and…

Portfolio Management · Quantitative Finance 2016-11-07 M. Abeille , E. Serie , A. Lazaric , X. Brokmann

The Linear Quadratic Gaussian (LQG) problem is a classic and widely studied model in optimal control, providing a fundamental framework for designing controllers for linear systems subject to process and observation noises. In recent years,…

Optimization and Control · Mathematics 2026-03-17 Haoran Li , Xun Li , Yuan-Hua Ni , Xuebo Zhang

Many applications -- including power systems, robotics, and economics -- involve a dynamical system interacting with a stochastic and hard-to-model environment. We adopt a reinforcement learning approach to control such systems.…

Optimization and Control · Mathematics 2025-08-26 Abed AlRahman Al Makdah , Oliver Kosut , Lalitha Sankar , Shaofeng Zou

The effectiveness of model-based versus model-free methods is a long-standing question in reinforcement learning (RL). Motivated by recent empirical success of RL on continuous control tasks, we study the sample complexity of popular…

Machine Learning · Computer Science 2019-02-05 Stephen Tu , Benjamin Recht

We consider infinite-horizon discounted Markov decision processes and study the convergence rates of the natural policy gradient (NPG) and the Q-NPG methods with the log-linear policy class. Using the compatible function approximation…

Machine Learning · Computer Science 2023-02-22 Rui Yuan , Simon S. Du , Robert M. Gower , Alessandro Lazaric , Lin Xiao

In this paper, we study the problem of stabilizing continuous-time switched linear systems with quantized output feedback. We assume that the observer and the control gain are given for each mode. Also, the plant mode is known to the…

Systems and Control · Computer Science 2015-09-03 Masashi Wakaiki , Yutaka Yamamoto

We provide a solution to the heretofore open problem of stabilization of systems with arbitrarily long delays at the input and output of a nonlinear system using output feedback only. The solution is global, employs the predictor approach…

Optimization and Control · Mathematics 2013-08-15 Iasson Karafyllis , Miroslav Krstic

In this paper, we propose a novel reinforcement- learning algorithm consisting in a stochastic variance-reduced version of policy gradient for solving Markov Decision Processes (MDPs). Stochastic variance-reduced gradient (SVRG) methods…

Machine Learning · Computer Science 2018-06-15 Matteo Papini , Damiano Binaghi , Giuseppe Canonaco , Matteo Pirotta , Marcello Restelli

Policy gradient methods are known to be highly sensitive to the choice of policy parameterization. In particular, the widely used softmax parameterization can induce ill-conditioned optimization landscapes and lead to exponentially slow…

Machine Learning · Computer Science 2026-04-02 Safwan Labbi , Daniil Tiapkin , Paul Mangold , Eric Moulines

Relative temporal-difference (TD) learning was introduced to mitigate the slow convergence of TD methods when the discount factor approaches one by subtracting a baseline from the temporal-difference update. While this idea has been studied…

Machine Learning · Computer Science 2026-04-08 Masoud S. Sakha , Rushikesh Kamalapurkar , Sean Meyn

We present new policy mirror descent (PMD) methods for solving reinforcement learning (RL) problems with either strongly convex or general convex regularizers. By exploring the structural properties of these overall highly nonconvex…

Machine Learning · Computer Science 2022-04-08 Guanghui Lan

The Sequential Linear Quadratic (SLQ) algorithm is a continuous-time variant of the well-known Differential Dynamic Programming (DDP) technique with a Gauss-Newton Hessian approximation. This family of methods has gained popularity in the…

Robotics · Computer Science 2021-03-29 Jean-Pierre Sleiman , Farbod Farshidian , Marco Hutter

In this paper, a polynomial chaos based framework for designing controllers for discrete time linear systems with probabilistic parameters is presented. Conditions for exponential-mean-square stability for such systems are derived and…

Optimization and Control · Mathematics 2020-04-06 Vaishnav Tadiparthi , Raktim Bhattacharya

We present a multi-query recovery policy for a hybrid system with goal limit cycle. The sample trajectories and the hybrid limit cycle of the dynamical system are stabilized using locally valid Time Varying LQR controller policies which…

Robotics · Computer Science 2017-11-15 Ramkumar Natarajan , Siddharthan Rajasekaran , Jonathan D. Taylor

Flow $Q$-learning has recently been introduced to integrate learning from expert demonstrations into an actor-critic structure. Central to this innovation is the ``the one-step policy'' network, which is optimized through a $Q$-function…

Systems and Control · Electrical Eng. & Systems 2025-11-17 Farnaz Adib Yaghmaie , Arunava Naha

This paper introduces and analyzes an improved Q-learning algorithm for discrete-time linear time-invariant systems. The proposed method does not require any knowledge of the system dynamics, and it enjoys significant efficiency advantages…

Systems and Control · Electrical Eng. & Systems 2023-04-03 Victor G. Lopez , Mohammad Alsalti , Matthias A. Müller

Recent advances in rule-based reinforcement learning (RL) have significantly improved the reasoning capability of language models (LMs) with rule-based rewards. However, existing RL methods -- such as GRPO, REINFORCE++, and RLOO -- often…

Machine Learning · Computer Science 2025-05-20 Zongkai Liu , Fanqing Meng , Lingxiao Du , Zhixiang Zhou , Chao Yu , Wenqi Shao , Qiaosheng Zhang

We study the problem of stabilizing an unknown partially observable linear time-invariant (LTI) system. For fully observable systems, leveraging an unstable/stable subspace decomposition approach, state-of-art sample complexity is…

Systems and Control · Electrical Eng. & Systems 2025-03-24 Ziyi Zhang , Yorie Nakahira , Guannan Qu

This paper presents a novel approach to reinforcement learning (RL) for control systems that provides probabilistic stability guarantees using finite data. Leveraging Lyapunov's method, we propose a probabilistic stability theorem that…

Machine Learning · Computer Science 2026-03-03 Minghao Han , Lixian Zhang , Chenliang Liu , Zhipeng Zhou , Jun Wang , Wei Pan

This work proposes a two-layered control scheme for constrained nonlinear systems represented by a class of recurrent neural networks and affected by additive disturbances. In particular, a base controller ensures global or regional…

Systems and Control · Electrical Eng. & Systems 2026-03-27 Daniele Ravasio , Danilo Saccani , Marcello Farina , Giancarlo Ferrari-Trecate