Related papers: Policy Iteration for Exploratory Hamilton--Jacobi-…
Despite its popularity in the reinforcement learning community, a provably convergent policy gradient method for continuous space-time control problems with nonlinear state dynamics has been elusive. This paper proposes proximal gradient…
Reflected diffusions naturally arise in many problems from applications ranging from economics and mathematical biology to queueing theory. In this paper we consider a class of infinite time-horizon singular stochastic control problems for…
Decision-making problems in uncertain or stochastic domains are often formulated as Markov decision processes (MDPs). Policy iteration (PI) is a popular algorithm for searching over policy-space, the size of which is exponential in the…
Convex Q-learning is a recent approach to reinforcement learning, motivated by the possibility of a firmer theory for convergence, and the possibility of making use of greater a priori knowledge regarding policy or value function structure.…
We prove convergence of the proximal policy gradient method for a class of constrained stochastic control problems with control in both the drift and diffusion of the state process. The problem requires either the running or terminal cost…
The scaling invariance for chaotic orbits near a transition from unlimited to limited diffusion in a dissipative standard mapping is explained via the analytical solution of the diffusion equation. It gives the probability of observing a…
The framework of deep operator network (DeepONet) has been widely exploited thanks to its capability of solving high dimensional partial differential equations. In this paper, we incorporate DeepONet with a recently developed policy…
We consider discrete-time infinite horizon deterministic optimal control problems with nonnegative cost per stage, and a destination that is cost-free and absorbing. The classical linear-quadratic regulator problem is a special case. Our…
In this paper, which is a continuation of the previously published discrete time paper we develop a theory for continuous time stochastic control problems which, in various ways, are time inconsistent in the sense that they do not admit a…
We study a continuous-time, finite horizon, stochastic partially reversible investment problem for a firm producing a single good in a market with frictions. The production capacity is modeled as a one-dimensional, time-homogeneous, linear…
The path-integral control, which stems from the stochastic Hamilton-Jacobi-Bellman equation, is one of the methods to control stochastic nonlinear systems. This paper gives a new insight into nonlinear stochastic optimal control problems…
In this paper we propose an on-line policy iteration (PI) algorithm for finite-state infinite horizon discounted dynamic programming, whereby the policy improvement operation is done on-line, only for the states that are encountered during…
The Hamilton-Jacobi-Bellman equation (HJB) associated with the time inhomogeneous singular control problem is a parabolic partial differential equation, and the existence of a classical solution is usually difficult to prove. In this paper,…
This article studies a portfolio optimization problem, where the market consisting of several stocks is modeled by a multi-dimensional jump-diffusion process with age-dependent semi-Markov modulated coefficients. We study risk sensitive…
We introduce some approximation schemes for linear and fully non-linear diffusion equations of Bellman-Isaacs type. Although they are not monotone one can prove their convergence to the viscosity solution of the problem. Effective…
We examine the problem of two-point boundary optimal control of nonlinear systems over finite-horizon time periods with unknown model dynamics by employing reinforcement learning. We use techniques from singular perturbation theory to…
This paper deals with a class of neural SDEs and studies the limiting behavior of the associated sampled optimal control problems as the sample size grows to infinity. The neural SDEs with $N$ samples can be linked to the $N$-particle…
We consider a general class of stochastic optimal control problems, where the state process lives in a real separable Hilbert space and is driven by a cylindrical Brownian motion and a Poisson random measure; no special structure is imposed…
In this paper, we study distributional reinforcement learning from the perspective of statistical efficiency. We investigate distributional policy evaluation, aiming to estimate the complete return distribution (denoted $\eta^\pi$) attained…
Controlling systems of ordinary differential equations (ODEs) is ubiquitous in science and engineering. For finding an optimal feedback controller, the value function and associated fundamental equations such as the Bellman equation and the…