Related papers: How are policy gradient methods affected by the li…
In data-driven control, a central question is how to handle noisy data. In this work, we consider the problem of designing a stabilizing controller for an unknown linear system using only a finite set of noisy data collected from the…
We prove convergence of the proximal policy gradient method for a class of constrained stochastic control problems with control in both the drift and diffusion of the state process. The problem requires either the running or terminal cost…
Stochastic dynamical systems allow modelling of transitions induced by disturbances, in particular from an attracting equilibrium and crossing the stable manifold of a saddle. In the small-noise limit, the probability of such transitions is…
In neural networks with binary activations and or binary weights the training by gradient descent is complicated as the model has piecewise constant response. We consider stochastic binary networks, obtained by adding noises in front of…
Off-policy stochastic actor-critic methods rely on approximating the stochastic policy gradient in order to derive an optimal policy. One may also derive the optimal policy by approximating the action-value gradient. The use of action-value…
We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-quadratic systems, and study various settings of driving…
Based on the heuristics that maintaining presumptions can be beneficial in uncertain environments, we propose a set of basic axioms for learning systems to incorporate the concept of prejudice. The simplest, memoryless model of a…
We extend recent analyses of stochastic effects in game dynamical learning to cases of multi-player games, and to games defined on networked structures. By means of an expansion in the noise strength we consider the weak-noise limit, and…
Dynamic oracles provide strong supervision for training constituency parsers with exploration, but must be custom defined for a given parser's transition system. We explore using a policy gradient method as a parser-agnostic alternative. In…
Diverse complex dynamical systems are known to exhibit abrupt regime shifts at bifurcation points of the saddle-node type. The dynamics of most of these systems, however, have a stochastic component resulting in noise driven regime shifts…
Model-free and model-based reinforcement learning are two ends of a spectrum. Learning a good policy without a dynamic model can be prohibitively expensive. Learning the dynamic model of a system can reduce the cost of learning the policy,…
Prediction via deterministic continuous-time models will always be subject to model error, for example due to unexplainable phenomena, uncertainties in any data driving the model, or discretisation/resolution issues. In this paper, we build…
Randomized experiments are the gold standard for evaluating the effects of changes to real-world systems. Data in these tests may be difficult to collect and outcomes may have high variance, resulting in potentially large measurement error.…
We study the stationary states of variants of the noisy voter model, subject to fluctuating parameters or external environments. Specifically, we consider scenarios in which the herding-to-noise ratio switches randomly and on different time…
Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet expressive setting for analyzing the optimization and…
Gene expression is inherently noisy as many steps in the read-out of the genetic information are stochastic. To disentangle the effect of different sources of stochasticity in such systems, we consider various models that describe some…
Understanding the limitations of gradient methods, and stochastic gradient descent (SGD) in particular, is a central challenge in learning theory. To that end, a commonly used tool is the Statistical Queries (SQ) framework, which studies…
We study the noisy voter model using a specific non-linear dependence of the rates that takes into account collective interaction between individuals. The resulting model is solved exactly under the all-to-all coupling configuration and…
Nearly-elastic model systems with one or two degrees of freedom are considered: the system is undergoing a small loss of energy in each collision with the "wall". We show that instabilities in this purely deterministic system lead to…
Natural and formal languages provide an effective mechanism for humans to specify instructions and reward functions. We investigate how to generate policies via RL when reward functions are specified in a symbolic language captured by…