English
Related papers

Related papers: How are policy gradient methods affected by the li…

200 papers

Despite its popularity in the reinforcement learning community, a provably convergent policy gradient method for continuous space-time control problems with nonlinear state dynamics has been elusive. This paper proposes proximal gradient…

Optimization and Control · Mathematics 2022-12-27 Christoph Reisinger , Wolfgang Stockinger , Yufei Zhang

Policy gradients in continuous control have been derived for both stochastic and deterministic policies. Here we study the relationship between the two. In a widely-used family of MDPs involving Gaussian control noise and quadratic control…

Machine Learning · Computer Science 2025-06-11 Emo Todorov

This paper is concerned with the problem of Model Predictive Control and Rolling Horizon Control of discrete-time systems subject to possibly unbounded random noise inputs, while satisfying hard bounds on the control inputs. We use a…

Optimization and Control · Mathematics 2010-09-08 Peter Hokayem , Debasish Chatterjee , John Lygeros

Policy gradient methods have enabled deep reinforcement learning (RL) to approach challenging continuous control problems, even when the underlying systems involve highly nonlinear dynamics that generate complex non-smooth optimization…

Machine Learning · Computer Science 2024-05-29 Tao Wang , Sylvia Herbert , Sicun Gao

The multidimensional Uncertain Volatility Model leads to robust option pricing problems under joint volatility and correlation uncertainty. Their numerical resolution quickly becomes challenging because the associated stochastic control…

Computational Finance · Quantitative Finance 2026-05-11 Lokman A Abbas-Turki , Jean-François Chassagneux , Jean-Philippe Lemor , Grégoire Loeper , Simon Sananes

In this paper, we investigate constrained control of continuous-time linear stochastic systems. We show that for certain system parameter settings, constrained control policies can never achieve stabilization. Specifically, we explore a…

Systems and Control · Electrical Eng. & Systems 2021-08-17 Ahmet Cetinkaya , Masako Kishida

Policy-gradient methods are widely used in reinforcement learning, yet training often becomes unstable or slows down as learning progresses. We study this phenomenon through the noise-to-signal ratio (NSR) of a policy-gradient estimator,…

Optimization and Control · Mathematics 2026-02-10 Haoyu Han , Heng Yang

Recent work on data-driven control and reinforcement learning has renewed interest in a relative old field in control theory: model-free optimal control approaches which work directly with a cost function and do not rely upon perfect…

Optimization and Control · Mathematics 2021-08-31 Eduardo D. Sontag

Policy gradients methods apply to complex, poorly understood, control problems by performing stochastic gradient descent over a parameterized class of polices. Unfortunately, even for simple control problems solvable by standard dynamic…

Machine Learning · Computer Science 2022-06-22 Jalaj Bhandari , Daniel Russo

We revisit the finite time analysis of policy gradient methods in the one of the simplest settings: finite state and action MDPs with a policy class consisting of all stochastic policies and with exact gradient evaluations. There has been…

Machine Learning · Computer Science 2021-12-14 Jalaj Bhandari , Daniel Russo

Policy gradient (PG) methods are successful approaches to deal with continuous reinforcement learning (RL) problems. They learn stochastic parametric (hyper)policies by either exploring in the space of actions or in the space of parameters.…

Machine Learning · Computer Science 2024-05-31 Alessandro Montenegro , Marco Mussi , Alberto Maria Metelli , Matteo Papini

Policy gradient methods have been frequently applied to problems in control and reinforcement learning with great success, yet existing convergence analysis still relies on non-intuitive, impractical and often opaque conditions. In…

Machine Learning · Computer Science 2022-04-08 Matthew S. Zhang , Murat A. Erdogdu , Animesh Garg

We study the limiting dynamics of a large class of noisy gradient descent systems in the overparameterized regime. In this regime the set of global minimizers of the loss is large, and when initialized in a neighbourhood of this zero-loss…

Machine Learning · Computer Science 2024-04-19 Anna Shalova , André Schlichting , Mark Peletier

We study a system whose dynamics are governed by predictions of its future states. A general formalism and concrete examples are presented. We find that the dynamical characteristics depend on how to shape the predictions as well as on how…

Other Condensed Matter · Physics 2015-06-25 Toru Ohira

In this paper, we examine the fundamental performance limitations in the control of stochastic dynamical systems; more specifically, we derive generic $\mathcal{L}_p$ bounds that hold for any causal (stabilizing) controllers and any…

Systems and Control · Electrical Eng. & Systems 2021-06-07 Song Fang , Quanyan Zhu

Political moderation, a key attractor in democratic systems, proves highly fragile under realistic information conditions. We develop a stochastic model of opinion dynamics to analyze how noise and differential susceptibility reshape the…

Physics and Society · Physics 2026-01-23 Renato Vieira dos Santos

We consider the failure of localized control in a nonlinear spatially extended system caused by extremely small amounts of noise. It is shown that this failure occurs as a result of a nonlinear instability. Nonlinear instabilities can occur…

Pattern Formation and Solitons · Physics 2009-11-07 Roman O. Grigoriev , Andreas Handel

The asymptotic behavior of the stochastic gradient algorithm with a biased gradient estimator is analyzed. Relying on arguments based on the dynamic system theory (chain-recurrence) and the differential geometry (Yomdin theorem and…

Statistics Theory · Mathematics 2017-09-04 Vladislav B. Tadic , Arnaud Doucet

Reinforcement learning is a promising approach to learning robotics controllers. It has recently been shown that algorithms based on finite-difference estimates of the policy gradient are competitive with algorithms based on the policy…

Machine Learning · Computer Science 2021-10-12 Osbert Bastani

We consider a data-driven formulation of the classical discrete-time stochastic control problem. Our approach exploits the natural structure of many such problems, in which significant portions of the system are uncontrolled. Employing the…

Optimization and Control · Mathematics 2025-08-25 Boris Baros , Samuel N. Cohen , Christoph Reisinger
‹ Prev 1 2 3 10 Next ›