Related papers: Policy Iteration for Exploratory Hamilton--Jacobi-…
We study reinforcement learning for controlled diffusion processes with unbounded continuous state spaces, bounded continuous actions, and polynomially growing rewards: settings that arise naturally in finance, economics, and operations…
We study a class of singular stochastic control problems for a one-dimensional diffusion $X$ in which the performance criterion to be optimised depends explicitly on the running infimum $I$ (or supremum $S$) of the controlled process. We…
Folklore says that Howard's Policy Improvement Algorithm converges extraordinarily fast, even for controlled diffusion settings. In a previous paper, we proved that approximations of the solution of a particular parabolic partial…
We explore the approximation of feedback control of integro-differential equations containing a fractional Laplacian term. To obtain feedback control for the state variable of this nonlocal equation we use the Hamilton--Jacobi--Bellman…
This paper presents a {\delta}-PI algorithm which is based on damped Newton method for the H{\infty} tracking control problem of unknown continuous-time nonlinear system. A discounted performance function and an augmented system are used to…
We study the global linear convergence of policy gradient (PG) methods for finite-horizon continuous-time exploratory linear-quadratic control (LQC) problems. The setting includes stochastic LQC problems with indefinite costs and allows…
Policy iteration is one of the classical frameworks of reinforcement learning, which requires a known initial stabilizing control. However, finding the initial stabilizing control depends on the known system model. To relax this requirement…
In this paper, we aim to develop the theory of optimal stochastic control for branching diffusion processes where both the movement and the reproduction of the particles depend on the control. More precisely, we study the problem of…
Bounded policy iteration is an approach to solving infinite-horizon POMDPs that represents policies as stochastic finite-state controllers and iteratively improves a controller by adjusting the parameters of each node using linear…
We study the problem of learning the optimal control policy for fine-tuning a given diffusion process, using general value function approximation. We develop a new class of algorithms by solving a variational inequality problem based on the…
In this paper, we study the optimal dividend problem under the continuous time diffusion model with the bounded dividend rate from the Reinforcement Learning (RL) perspective. Unlike the standard literature, our main focus will be on…
Continuous-time stochastic processes underlie many natural and engineered systems. In healthcare, autonomous driving, and industrial control, direct interaction with the environment is often unsafe or impractical, motivating offline…
An optimal control problem is considered for a stochastic differential equation with the cost functional determined by a backward stochastic Volterra integral equation (BSVIE, for short). This kind of cost functional can cover the general…
This paper considers time-inconsistent problems when control and stopping strategies are required to be made simultaneously (called stopping control problems by us). We first formulate the timeinconsistent stopping control problems under…
This paper is concerned with the axiomatic foundation and explicit construction of a general class of optimality criteria that can be used for investment problems with multiple time horizons, or when the time horizon is not known in…
Following the recent resurgence in establishing linear control theoretic benchmarks for reinforcement leaning (RL)-based policy optimization (PO) for complex dynamical systems with continuous state and action spaces, an optimal control…
This paper employs a policy iteration reinforcement learning (RL) method to study continuous-time linear-quadratic mean-field control problems in infinite horizon. The drift and diffusion terms in the dynamics involve the states, the…
We establish a connection between stochastic optimal control and generative models based on stochastic differential equations (SDEs), such as recently developed diffusion probabilistic models. In particular, we derive a…
In this paper, we propose a generalized successive approximation method (SAM), called invariantly admissible policy iteration (PI), for finding the solution to a class of input-affine nonlinear optimal control problems by iterations. Unlike…
We study stochastic optimal control problems for (possibly degenerate) McKean-Vlasov controlled diffusions and obtain discrete-time as well as finite interacting particle approximations. (i) Under mild assumptions, we first prove the…