English
Related papers

Related papers: How are policy gradient methods affected by the li…

200 papers

In data-driven control, a central question is how to handle noisy data. In this work, we consider the problem of designing a stabilizing controller for an unknown linear system using only a finite set of noisy data collected from the…

Systems and Control · Electrical Eng. & Systems 2021-06-29 Andrea Bisoffi , Claudio De Persis , Pietro Tesi

We prove convergence of the proximal policy gradient method for a class of constrained stochastic control problems with control in both the drift and diffusion of the state process. The problem requires either the running or terminal cost…

Optimization and Control · Mathematics 2025-05-27 Ashley Davey , Harry Zheng

Stochastic dynamical systems allow modelling of transitions induced by disturbances, in particular from an attracting equilibrium and crossing the stable manifold of a saddle. In the small-noise limit, the probability of such transitions is…

Statistical Mechanics · Physics 2025-09-05 Jiayao Shao , Tobias Grafke , Robert S. MacKay

In neural networks with binary activations and or binary weights the training by gradient descent is complicated as the model has piecewise constant response. We consider stochastic binary networks, obtained by adding noises in front of…

Machine Learning · Statistics 2020-11-05 Alexander Shekhovtsov , Viktor Yanush , Boris Flach

Off-policy stochastic actor-critic methods rely on approximating the stochastic policy gradient in order to derive an optimal policy. One may also derive the optimal policy by approximating the action-value gradient. The use of action-value…

Machine Learning · Statistics 2017-03-14 Yemi Okesanjo , Victor Kofia

We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-quadratic systems, and study various settings of driving…

Machine Learning · Computer Science 2020-05-19 Dhruv Malik , Ashwin Pananjady , Kush Bhatia , Koulik Khamaru , Peter L. Bartlett , Martin J. Wainwright

Based on the heuristics that maintaining presumptions can be beneficial in uncertain environments, we propose a set of basic axioms for learning systems to incorporate the concept of prejudice. The simplest, memoryless model of a…

Adaptation and Self-Organizing Systems · Physics 2007-05-23 Andreas U. Schmidt

We extend recent analyses of stochastic effects in game dynamical learning to cases of multi-player games, and to games defined on networked structures. By means of an expansion in the noise strength we consider the weak-noise limit, and…

Physics and Society · Physics 2012-04-20 Alex J. Bladon , Tobias Galla

Dynamic oracles provide strong supervision for training constituency parsers with exploration, but must be custom defined for a given parser's transition system. We explore using a policy gradient method as a parser-agnostic alternative. In…

Computation and Language · Computer Science 2018-06-11 Daniel Fried , Dan Klein

Diverse complex dynamical systems are known to exhibit abrupt regime shifts at bifurcation points of the saddle-node type. The dynamics of most of these systems, however, have a stochastic component resulting in noise driven regime shifts…

Statistical Mechanics · Physics 2013-10-29 Sayantari Ghosh , Amit Kumar Pal , Indrani Bose

Model-free and model-based reinforcement learning are two ends of a spectrum. Learning a good policy without a dynamic model can be prohibitively expensive. Learning the dynamic model of a system can reduce the cost of learning the policy,…

Robotics · Computer Science 2022-01-19 Arash Mehrjou , Ashkan Soleymani , Stefan Bauer , Bernhard Schölkopf

Prediction via deterministic continuous-time models will always be subject to model error, for example due to unexplainable phenomena, uncertainties in any data driving the model, or discretisation/resolution issues. In this paper, we build…

Dynamical Systems · Mathematics 2025-06-30 Liam Blake , John Maclean , Sanjeeva Balasuriya

Randomized experiments are the gold standard for evaluating the effects of changes to real-world systems. Data in these tests may be difficult to collect and outcomes may have high variance, resulting in potentially large measurement error.…

Machine Learning · Statistics 2018-06-27 Benjamin Letham , Brian Karrer , Guilherme Ottoni , Eytan Bakshy

We study the stationary states of variants of the noisy voter model, subject to fluctuating parameters or external environments. Specifically, we consider scenarios in which the herding-to-noise ratio switches randomly and on different time…

Physics and Society · Physics 2023-06-01 Annalisa Caligiuri , Tobias Galla

Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet expressive setting for analyzing the optimization and…

Optimization and Control · Mathematics 2026-05-19 Begoña García Malaxechebarría , Courtney Paquette , Maryam Fazel , Dmitriy Drusvyatskiy

Gene expression is inherently noisy as many steps in the read-out of the genetic information are stochastic. To disentangle the effect of different sources of stochasticity in such systems, we consider various models that describe some…

Molecular Networks · Quantitative Biology 2015-06-05 Rahul Marathe , David Gomez , Stefan Klumpp

Understanding the limitations of gradient methods, and stochastic gradient descent (SGD) in particular, is a central challenge in learning theory. To that end, a commonly used tool is the Statistical Queries (SQ) framework, which studies…

Machine Learning · Computer Science 2026-02-06 Daniel Barzilai , Ohad Shamir

We study the noisy voter model using a specific non-linear dependence of the rates that takes into account collective interaction between individuals. The resulting model is solved exactly under the all-to-all coupling configuration and…

Physics and Society · Physics 2018-10-05 A. F. Peralta , A. Carro , M. San Miguel , R. Toral

Nearly-elastic model systems with one or two degrees of freedom are considered: the system is undergoing a small loss of energy in each collision with the "wall". We show that instabilities in this purely deterministic system lead to…

Probability · Mathematics 2012-08-31 Mark Freidlin , Wenqing Hu

Natural and formal languages provide an effective mechanism for humans to specify instructions and reward functions. We investigate how to generate policies via RL when reward functions are specified in a symbolic language captured by…

Machine Learning · Computer Science 2022-11-24 Andrew C. Li , Zizhao Chen , Pashootan Vaezipoor , Toryn Q. Klassen , Rodrigo Toro Icarte , Sheila A. McIlraith