Related papers: Bootstrap Policy Iteration for Stochastic LQ Track…
A general and new stochastic linear quadratic optimal control problem is studied, where the coefficients are allowed to be time-varying, and both state delay and control delay can appear simultaneously in the state equation and the cost…
A linear-quadratic optimal control problem for a forward stochastic Volterra integral equation (FSVIE, for short) is considered. Under the usual convexity conditions, open-loop optimal control exists, which can be characterized by the…
This paper applies a reinforcement learning (RL) method to solve infinite horizon continuous-time stochastic linear quadratic problems, where drift and diffusion terms in the dynamics may depend on both the state and control. Based on…
We consider the problem of nonlinear stochastic optimal control. This problem is thought to be fundamentally intractable owing to Bellman's "curse of dimensionality". We present a result that shows that repeatedly solving an open-loop…
This paper examines learning the optimal filtering policy, known as the Kalman gain, for a linear system with unknown noise covariance matrices using noisy output data. The learning problem is formulated as a stochastic policy optimization…
In this paper, a leader-follower stochastic differential game is studied for a linear stochastic differential equation with a quadratic cost functional. The coefficients in the state equation and the weighting matrices in the cost…
The optimal disturbance rejection control problem is considered for consensus tracking systems affected by external persistent disturbances and noise. Optimal estimated values of system states are obtained by recursive filtering for the…
Linear dynamical systems that obey stochastic differential equations are canonical models. While optimal control of known systems has a rich literature, the problem is technically hard under model uncertainty and there are hardly any…
The fundamental lemma by Jan C. Willems and co-authors enables the representation of all input-output trajectories of a linear time-invariant system by measured input-output data. This result has proven to be pivotal for data-driven…
In this paper,we mainly focus on the numerical solution of high-dimensional stochastic optimal control problem driven by fully-coupled forward-backward stochastic differential equations (FBSDEs in short) through deep learning. We first…
This paper presents a novel synthesis method for designing an optimal and robust guidance law for a non-throttleable upper stage of a launch vehicle, using a convex approach. In the unperturbed scenario, a combination of lossless and…
This article studies the control ideas of the optimal backstepping technique, proposing an event-triggered optimal tracking control scheme for a class of strict-feedback nonlinear systems with non-affine and nonlinear faults. A simplified…
We are motivated by the real challenges presented in a human-robot system to develop new designs that are efficient at data level and with performance guarantees such as stability and optimality at systems level. Existing…
We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework…
Policy gradient (PG) methods are the backbone of many reinforcement learning algorithms due to their good performance in policy optimization problems. As a gradient-based approach, PG methods typically rely on knowledge of the system…
This paper addresses the problem of model-free reinforcement learning for Robust Markov Decision Process (RMDP) with large state spaces. The goal of the RMDP framework is to find a policy that is robust against the parameter uncertainties…
This article proposes an improved trajectory optimization approach for stochastic optimal control of dynamical systems affected by measurement noise by combining optimal control with maximum likelihood techniques to improve the reduction of…
The task of dialog management is commonly decomposed into two sequential subtasks: dialog state tracking and dialog policy learning. In an end-to-end dialog system, the aim of dialog state tracking is to accurately estimate the true dialog…
We provide a technique to obtain provably optimal control sequences for quantum systems under the influence of time-correlated multiplicative control noise. Utilizing the circuit-level noise model introduced in [Phys. Rev. Research 3,…
Many of the recent trajectory optimization algorithms alternate between linear approximation of the system dynamics around the mean trajectory and conservative policy update. One way of constraining the policy change is by bounding the…