English
Related papers

Related papers: How are policy gradient methods affected by the li…

200 papers

Policy robustness in Reinforcement Learning may not be desirable at any cost: the alterations caused by robustness requirements from otherwise optimal policies should be explainable, quantifiable and formally verifiable. In this work we…

Machine Learning · Computer Science 2023-12-12 Daniel Jarne Ornia , Licio Romao , Lewis Hammond , Manuel Mazo , Alessandro Abate

We design receding horizon control strategies for stochastic discrete-time linear systems with additive (possibly) unbounded disturbances, while obeying hard bounds on the control inputs. We pose the problem of selecting an appropriate…

Optimization and Control · Mathematics 2011-07-07 Debasish Chatterjee , Peter Hokayem , John Lygeros

Considering generating samples with high rewards, we focus on optimizing deep neural networks parameterized stochastic differential equations (SDEs), the advanced generative models with high expressiveness, with policy gradient, the leading…

Machine Learning · Computer Science 2024-06-27 Xiangxin Zhou , Liang Wang , Yichi Zhou

We consider the influence of stochastic perturbations on stability of a unique positive equilibrium of a difference equation subject to prediction-based control. These perturbations may be multiplicative $$x_{n+1}=f(x_n)-\left( \alpha +…

Dynamical Systems · Mathematics 2016-06-08 Elena Braverman , Conall Kelly , Alexandra Rodkina

We prove that stochastic gradient descent efficiently converges to the global optimizer of the maximum likelihood objective of an unknown linear time-invariant dynamical system from a sequence of noisy observations generated by the system.…

Machine Learning · Computer Science 2019-02-12 Moritz Hardt , Tengyu Ma , Benjamin Recht

We study the variance of the REINFORCE policy gradient estimator in environments with continuous state and action spaces, linear dynamics, quadratic cost, and Gaussian noise. These simple environments allow us to derive bounds on the…

Machine Learning · Computer Science 2019-10-04 James A. Preiss , Sébastien M. R. Arnold , Chen-Yu Wei , Marius Kloft

Policy gradient lies at the core of deep reinforcement learning (RL) in continuous domains. Despite much success, it is often observed in practice that RL training with policy gradient can fail for many reasons, even on standard control…

Machine Learning · Computer Science 2024-01-23 Tao Wang , Sylvia Herbert , Sicun Gao

This paper considers the problem of learning safe policies in the context of reinforcement learning (RL). In particular, we consider the notion of probabilistic safety. This is, we aim to design policies that maintain the state of the…

Machine Learning · Computer Science 2023-04-20 Weiqin Chen , Dharmashankar Subramanian , Santiago Paternain

We present a unified framework for learning continuous control policies using backpropagation. It supports stochastic control by treating stochasticity in the Bellman equation as a deterministic function of exogenous noise. The product is a…

Machine Learning · Computer Science 2015-11-02 Nicolas Heess , Greg Wayne , David Silver , Timothy Lillicrap , Yuval Tassa , Tom Erez

In this paper, a non-autonomous stochastic logistic system is considered. An interesting result on the effect of stochastically perturbation for the dynamic behavior are obtained. That is, under certain conditions the stochastic system have…

Dynamical Systems · Mathematics 2012-08-08 Hu Hongxiao

Policy gradient algorithms are widely used in reinforcement learning and belong to the class of approximate dynamic programming methods. This paper studies two key policy gradient algorithms, the Natural Policy Gradient and the Gauss-Newton…

Systems and Control · Electrical Eng. & Systems 2026-05-11 Bowen Song , Sebastien Gros , Andrea Iannelli

We consider an agent trying to bring a system to an acceptable state by repeated probabilistic action. Several recent works on algorithmizations of the Lovasz Local Lemma (LLL) can be seen as establishing sufficient conditions for the agent…

Discrete Mathematics · Computer Science 2016-11-29 Dimitris Achlioptas , Fotis Iliopoulos , Nikos Vlassis

We study the problem of system identification for stochastic continuous-time dynamics, based on a single finite-length state trajectory. We present a method for estimating the possibly unstable open-loop matrix by employing properly…

Machine Learning · Statistics 2025-09-30 Reza Sadeghi Hafshejani , Mohamad Kazem Shirani Fradonbeh

We consider policy gradient methods for stochastic optimal control problem in continuous time. In particular, we analyze the gradient flow for the control, viewed as a continuous time limit of the policy gradient method. We prove the global…

Optimization and Control · Mathematics 2025-04-15 Mo Zhou , Jianfeng Lu

Fluctuations and noise may alter the behavior of dynamical systems considerably. For example, oscillations may be sustained by demographic fluctuations in biological systems where a stable fixed point is found in the absence of noise. We…

Adaptation and Self-Organizing Systems · Physics 2009-11-13 Richard P. Boland , Tobias Galla , Alan J. McKane

Despite the celebrated success of stochastic control approaches for uncertain systems, such approaches are limited in the ability to handle non-Gaussian uncertainties. This work presents an adaptive robust control for linear uncertain…

Optimization and Control · Mathematics 2026-01-13 Xuehui Ma , Shiliang Zhang , Zhiyong Sun , Xiaohui Zhang , Sabita Maharjan

Policy gradient methods are a widely used class of model-free reinforcement learning algorithms where a state-dependent baseline is used to reduce gradient estimator variance. Several recent papers extend the baseline to depend on both the…

Machine Learning · Computer Science 2018-11-20 George Tucker , Surya Bhupatiraju , Shixiang Gu , Richard E. Turner , Zoubin Ghahramani , Sergey Levine

The aim of the present paper is to provide necessary and sufficient conditions to maintain a stochastic coupled system, with porous media components and gradient-type noise in a prescribed set of constraints by using internal controls. This…

Analysis of PDEs · Mathematics 2022-02-08 Ioana Ciotir , Dan Goreac , Ionut Munteanu

A key limitation in using various modern methods of machine learning in developing feedback control policies is the lack of appropriate methodologies to analyze their long-term dynamics, in terms of making any sort of guarantees (even…

Machine Learning · Computer Science 2021-06-17 Sean Gillen , Katie Byl

We study the sample complexity of policy gradient for log-growth control -- the problem of learning, from observed state transitions, a feedback gain that optimally stabilizes a scalar linear system driven through a multiplicative-noise…

Systems and Control · Electrical Eng. & Systems 2026-05-27 Qiuhua Pan , Yukai Shen , Liwei Zhang , Cailian Chen , Xinping Guan