Related papers: How are policy gradient methods affected by the li…
We consider the problem of nonlinear stochastic optimal control. This problem is thought to be fundamentally intractable owing to Bellman's "curse of dimensionality". We present a result that shows that repeatedly solving an open-loop…
In recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-feedback control,…
In this paper, we study the control of dynamical systems under temporal logic task specifications using gradient-based methods relying on quantitative measures that express the extent to which the tasks are satisfied. A class of controllers…
Model Predictive Control is an extremely effective control method for systems with input and state constraints. Model Predictive Control performance heavily depends on the accuracy of the open-loop prediction. For systems with uncertainty…
Noisy dynamical models are employed to describe a wide range of phenomena. Since exact modeling of these phenomena requires access to their microscopic dynamics, whose time scales are typically much shorter than the observable time scales,…
Policy gradient based reinforcement learning algorithms coupled with neural networks have shown success in learning complex policies in the model free continuous action space control setting. However, explicitly parameterized policies are…
Many theories are formulated as constrained systems. We provide a mechanism that explains the origin of physical states of a constrained system by a process of selection of noiseless subsystems when the system is coupled to an external…
Simple dynamical systems -- with a small number of degrees of freedom -- can behave in a complex manner due to the presence of chaos. Such systems are most often (idealized) limiting cases of more realistic situations. Isolating a small…
We consider a slow passage through a point of loss of stability. If the passage is sufficiently slow, the dynamics are controlled by additive random disturbances, even if they are extremely small. We derive expressions for the `exit value'…
We present a theoretical analysis of some popular adaptive Stochastic Gradient Descent (SGD) methods in the small learning rate regime. Using the stochastic modified equations framework introduced by Li et al., we derive effective…
We extend observability metrics based on the empirical observability Gramian from deterministic nonlinear systems to nonlinear stochastic systems in order to capture the impact of process noise on observability. We demonstrate that the…
The effect of small-amplitude noise on excitable systems with large time-scale separation is analyzed. It is found that small random perturbations of the fast excitatory variable result in the onset of a quasi-deterministic limit cycle…
We study the dynamics of fronts when both inertial effects and external fluctuations are taken into account. Stochastic fluctuations are introduced as multiplicative noise arising from a control parameter of the system. Contrary to the…
We consider slow-fast systems of differential equations, in which both the slow and fast variables are perturbed by noise. When the deterministic system admits a uniformly asymptotically stable slow manifold, we show that the sample paths…
A machine learning technique is proposed for quantifying uncertainty in power system dynamics with spatiotemporally correlated stochastic forcing. We learn one-dimensional linear partial differential equations for the probability density…
This paper is on learning the Kalman gain by policy optimization method. Firstly, we reformulate the finite-horizon Kalman filter as a policy optimization problem of the dual system. Secondly, we obtain the global linear convergence of…
This paper investigates a type of instability that is linked to the greedy policy improvement in approximated reinforcement learning. We show empirically that non-deterministic policy improvement can stabilize methods like LSPI by…
Off-policy learning refers to the problem of learning the value function of a way of behaving, or policy, while following a different policy. Gradient-based off-policy learning algorithms, such as GTD and TDC/GQ, converge even when using…
Intrinsic noise in objective function and derivatives evaluations may cause premature termination of optimization algorithms. Evaluation complexity bounds taking this situation into account are presented in the framework of a deterministic…
Mean field optimal control problems are a class of optimization problems that arise from optimal control when applied to the many body setting. In the noisy case one has a set of controllable stochastic processes and a cost function that is…