中文
相关论文

相关论文: How are policy gradient methods affected by the li…

200 篇论文

Policy robustness in Reinforcement Learning may not be desirable at any cost: the alterations caused by robustness requirements from otherwise optimal policies should be explainable, quantifiable and formally verifiable. In this work we…

机器学习 · 计算机科学 2023-12-12 Daniel Jarne Ornia , Licio Romao , Lewis Hammond , Manuel Mazo , Alessandro Abate

We design receding horizon control strategies for stochastic discrete-time linear systems with additive (possibly) unbounded disturbances, while obeying hard bounds on the control inputs. We pose the problem of selecting an appropriate…

最优化与控制 · 数学 2011-07-07 Debasish Chatterjee , Peter Hokayem , John Lygeros

Considering generating samples with high rewards, we focus on optimizing deep neural networks parameterized stochastic differential equations (SDEs), the advanced generative models with high expressiveness, with policy gradient, the leading…

机器学习 · 计算机科学 2024-06-27 Xiangxin Zhou , Liang Wang , Yichi Zhou

We consider the influence of stochastic perturbations on stability of a unique positive equilibrium of a difference equation subject to prediction-based control. These perturbations may be multiplicative $$x_{n+1}=f(x_n)-\left( \alpha +…

动力系统 · 数学 2016-06-08 Elena Braverman , Conall Kelly , Alexandra Rodkina

We prove that stochastic gradient descent efficiently converges to the global optimizer of the maximum likelihood objective of an unknown linear time-invariant dynamical system from a sequence of noisy observations generated by the system.…

机器学习 · 计算机科学 2019-02-12 Moritz Hardt , Tengyu Ma , Benjamin Recht

We study the variance of the REINFORCE policy gradient estimator in environments with continuous state and action spaces, linear dynamics, quadratic cost, and Gaussian noise. These simple environments allow us to derive bounds on the…

机器学习 · 计算机科学 2019-10-04 James A. Preiss , Sébastien M. R. Arnold , Chen-Yu Wei , Marius Kloft

Policy gradient lies at the core of deep reinforcement learning (RL) in continuous domains. Despite much success, it is often observed in practice that RL training with policy gradient can fail for many reasons, even on standard control…

机器学习 · 计算机科学 2024-01-23 Tao Wang , Sylvia Herbert , Sicun Gao

This paper considers the problem of learning safe policies in the context of reinforcement learning (RL). In particular, we consider the notion of probabilistic safety. This is, we aim to design policies that maintain the state of the…

机器学习 · 计算机科学 2023-04-20 Weiqin Chen , Dharmashankar Subramanian , Santiago Paternain

We present a unified framework for learning continuous control policies using backpropagation. It supports stochastic control by treating stochasticity in the Bellman equation as a deterministic function of exogenous noise. The product is a…

机器学习 · 计算机科学 2015-11-02 Nicolas Heess , Greg Wayne , David Silver , Timothy Lillicrap , Yuval Tassa , Tom Erez

In this paper, a non-autonomous stochastic logistic system is considered. An interesting result on the effect of stochastically perturbation for the dynamic behavior are obtained. That is, under certain conditions the stochastic system have…

动力系统 · 数学 2012-08-08 Hu Hongxiao

Policy gradient algorithms are widely used in reinforcement learning and belong to the class of approximate dynamic programming methods. This paper studies two key policy gradient algorithms, the Natural Policy Gradient and the Gauss-Newton…

系统与控制 · 电气工程与系统科学 2026-05-11 Bowen Song , Sebastien Gros , Andrea Iannelli

We consider an agent trying to bring a system to an acceptable state by repeated probabilistic action. Several recent works on algorithmizations of the Lovasz Local Lemma (LLL) can be seen as establishing sufficient conditions for the agent…

离散数学 · 计算机科学 2016-11-29 Dimitris Achlioptas , Fotis Iliopoulos , Nikos Vlassis

We study the problem of system identification for stochastic continuous-time dynamics, based on a single finite-length state trajectory. We present a method for estimating the possibly unstable open-loop matrix by employing properly…

机器学习 · 统计学 2025-09-30 Reza Sadeghi Hafshejani , Mohamad Kazem Shirani Fradonbeh

We consider policy gradient methods for stochastic optimal control problem in continuous time. In particular, we analyze the gradient flow for the control, viewed as a continuous time limit of the policy gradient method. We prove the global…

最优化与控制 · 数学 2025-04-15 Mo Zhou , Jianfeng Lu

Fluctuations and noise may alter the behavior of dynamical systems considerably. For example, oscillations may be sustained by demographic fluctuations in biological systems where a stable fixed point is found in the absence of noise. We…

适应与自组织系统 · 物理学 2009-11-13 Richard P. Boland , Tobias Galla , Alan J. McKane

Despite the celebrated success of stochastic control approaches for uncertain systems, such approaches are limited in the ability to handle non-Gaussian uncertainties. This work presents an adaptive robust control for linear uncertain…

最优化与控制 · 数学 2026-01-13 Xuehui Ma , Shiliang Zhang , Zhiyong Sun , Xiaohui Zhang , Sabita Maharjan

Policy gradient methods are a widely used class of model-free reinforcement learning algorithms where a state-dependent baseline is used to reduce gradient estimator variance. Several recent papers extend the baseline to depend on both the…

机器学习 · 计算机科学 2018-11-20 George Tucker , Surya Bhupatiraju , Shixiang Gu , Richard E. Turner , Zoubin Ghahramani , Sergey Levine

The aim of the present paper is to provide necessary and sufficient conditions to maintain a stochastic coupled system, with porous media components and gradient-type noise in a prescribed set of constraints by using internal controls. This…

偏微分方程分析 · 数学 2022-02-08 Ioana Ciotir , Dan Goreac , Ionut Munteanu

A key limitation in using various modern methods of machine learning in developing feedback control policies is the lack of appropriate methodologies to analyze their long-term dynamics, in terms of making any sort of guarantees (even…

机器学习 · 计算机科学 2021-06-17 Sean Gillen , Katie Byl

We study the sample complexity of policy gradient for log-growth control -- the problem of learning, from observed state transitions, a feedback gain that optimally stabilizes a scalar linear system driven through a multiplicative-noise…

系统与控制 · 电气工程与系统科学 2026-05-27 Qiuhua Pan , Yukai Shen , Liwei Zhang , Cailian Chen , Xinping Guan