中文
相关论文

相关论文: How are policy gradient methods affected by the li…

200 篇论文

We consider the problem of estimating unknown parameters in stochastic differential equations driven by colored noise, which we model as a sequence of Gaussian stationary processes with decreasing correlation time. We aim to infer…

数值分析 · 数学 2024-12-30 Grigorios A. Pavliotis , Sebastian Reich , Andrea Zanoni

Pulse stabilization of cycles with Prediction-Based Control including noise and stochastic stabilization of maps with multiple equilibrium points is analyzed for continuous but, generally, non-smooth maps. Sufficient conditions of global…

动力系统 · 数学 2022-08-19 Elena Braverman , Alexandra Rodkina

We study how the behavior of deep policy gradient algorithms reflects the conceptual framework motivating their development. To this end, we propose a fine-grained analysis of state-of-the-art methods based on key elements of this…

We investigate constrained optimal control problems for linear stochastic dynamical systems evolving in discrete time. We consider minimization of an expected value cost over a finite horizon. Hard constraints are introduced first, and then…

最优化与控制 · 数学 2011-07-07 Eugenio Cinquemani , Mayank Agarwal , Debasish Chatterjee , John Lygeros

We consider the stabilization of an unstable discrete-time linear system that is observed over a channel corrupted by continuous multiplicative noise. Our main result shows that if the system growth is large enough, then the system cannot…

系统与控制 · 计算机科学 2016-12-22 Jian Ding , Yuval Peres , Gireeja Ranade , Alex Zhai

In this paper, we revisit energy-based concepts of controllability and reformulate them for control-affine nonlinear systems perturbed by white noise. Specifically, we discuss the relation between controllability of deterministic systems…

最优化与控制 · 数学 2021-08-03 Carsten Hartmann , Lara Neureither , Markus Strehlau

In this work, we propose a stochastic gradient descent (SGD) framework to design data-driven policy gradient descent algorithms for the linear quadratic regulator problem. Two alternative schemes are considered to estimate the policy…

系统与控制 · 电气工程与系统科学 2026-02-24 Bowen Song , Simon Weissmann , Mathias Staudigl , Andrea Iannelli

Policy gradient is a generic and flexible reinforcement learning approach that generally enjoys simplicity in analysis, implementation, and deployment. In the last few decades, this approach has been extensively advanced for fully…

机器学习 · 计算机科学 2020-05-26 Kamyar Azizzadenesheli , Yisong Yue , Animashree Anandkumar

Deterministic models are approximations of reality that are easy to interpret and often easier to build than stochastic alternatives. Unfortunately, as nature is capricious, observational data can never be fully explained by deterministic…

机器学习 · 计算机科学 2020-03-31 Andrew Warrington , Saeid Naderiparizi , Frank Wood

Policy gradient methods are extensively used in reinforcement learning as a way to optimize expected return. In this paper, we explore the evolution of the policy parameters, for a special class of exactly solvable POMDPs, as a…

机器学习 · 计算机科学 2020-11-04 Gavin McCracken , Colin Daniels , Rosie Zhao , Anna Brandenberger , Prakash Panangaden , Doina Precup

We provide a new perspective to understand why reinforcement learning (RL) struggles with robustness and generalization. We show, by examples, that local optimal policies may contain unstable control for some dynamic parameters and…

机器人学 · 计算机科学 2023-11-03 Bing Song , Jean-Jacques Slotine , Quang-Cuong Pham

Policy gradient methods in reinforcement learning update policy parameters by taking steps in the direction of an estimated gradient of policy value. In this paper, we consider the statistically efficient estimation of policy gradients from…

机器学习 · 统计学 2020-02-21 Nathan Kallus , Masatoshi Uehara

Offline reinforcement learning, wherein one uses off-policy data logged by a fixed behavior policy to evaluate and learn new policies, is crucial in applications where experimentation is limited such as medicine. We study the estimation of…

机器学习 · 计算机科学 2020-06-09 Nathan Kallus , Masatoshi Uehara

People make strategic decisions many times a day - during negotiations, when coordinating actions with others, or when choosing partners for cooperation. The resulting dynamics can be studied with learning theory and evolutionary game…

种群与进化 · 定量生物学 2026-03-26 Marta C. Couto , Fernando P. Santos , Christian Hilbe

Stochastic processes offer a flexible mathematical formalism to model and reason about systems. Most analysis tools, however, start from the premises that models are fully specified, so that any parameters controlling the system's dynamics…

系统与控制 · 计算机科学 2017-01-11 Luca Bortolussi , Guido Sanguinetti

Machine learning models trained by different optimization algorithms under different data distributions can exhibit distinct generalization behaviors. In this paper, we analyze the generalization of models trained by noisy iterative…

机器学习 · 统计学 2022-12-29 Hao Wang , Rui Gao , Flavio P. Calmon

We consider infinite-horizon discounted Markov decision problems with finite state and action spaces and study the convergence rates of the projected policy gradient method and a general class of policy mirror descent methods, all with…

最优化与控制 · 数学 2022-03-08 Lin Xiao

Policy gradient methods are very attractive in reinforcement learning due to their model-free nature and convergence guarantees. These methods, however, suffer from high variance in gradient estimation, resulting in poor sample efficiency.…

机器学习 · 计算机科学 2018-11-16 Sergey Pankov

Policy-gradient methods have received increased attention recently as a mechanism for learning to act in partially observable environments. They have shown promise for problems admitting memoryless policies but have been less successful…

机器学习 · 计算机科学 2025-12-08 Douglas Aberdeen , Jonathan Baxter

We analyse the effect of intrinsic fluctuations on the properties of bistable stochastic systems with time scale separation operating under1 quasi-steady state conditions. We first formulate a stochastic generalisation of the quasi-steady…

生物物理 · 物理学 2016-03-02 Roberto de la Cruz , Pilar Guerrero , Fabian Spill , Tomás Alarcón