Related papers: Is Bellman Equation Enough for Learning Control?
We consider the problem of Reinforcement Learning for nonlinear stochastic dynamical systems. We show that in the RL setting, there is an inherent ``Curse of Variance" in addition to Bellman's infamous ``Curse of Dimensionality", in…
This paper is concerned with a stochastic recursive optimal control problem with time delay, where the controlled system is described by a stochastic differential delayed equation (SDDE) and the cost functional is formulated as the solution…
A Deterministic affine quadratic optimal control problem is considered. Due to the nature of the problem, optimal controls exist under some very mild conditions. Further, it is shown that under some assumptions, the value function is…
Reinforcement learning based adaptive/approximate dynamic programming (ADP) is a powerful technique to determine an approximate optimal controller for a dynamical system. These methods bypass the need to analytically solve the nonlinear…
In this article we study a finite horizon optimal control problem with monotone controls. We consider the associated Hamilton-Jacobi-Bellman (HJB) equation which characterizes the value function. We consider the totally discretized problem…
This note lays part of the theoretical ground for a definition of differential systems modeling reinforcement learning in continuous time non-Markovian rough environments. Specifically we focus on optimal relaxed control of rough equations…
This paper introduces the Hamilton-Jacobi-Bellman Proximal Policy Optimization (HJBPPO) algorithm into reinforcement learning. The Hamilton-Jacobi-Bellman (HJB) equation is used in control theory to evaluate the optimality of the value…
In this paper we study a class of stochastic control problems in which the control of the jump size is essential. Such a model is a generalized version for various applied problems ranging from optimal reinsurance selections for general…
We obtain weighted uniform estimates for the gradient of the solutions to a class of linear parabolic Cauchy problems with unbounded coefficients. Such estimates are then used to prove existence and uniqueness of the mild solution to a…
Using results from quantum filtering theory and methods from classical control theory, we derive an optimal control strategy for an open two-level system (a qubit in interaction with the electromagnetic field) controlled by a laser. The aim…
In this paper, we propose a novel image restoration framework that integrates optimal control techniques with the Hamilton-Jacobi-Bellman (HJB) equation. Motivated by models from production planning, our method restores degraded images by…
In this note, we discuss a class of time-dependent Hamilton-Jacobi equations depending on a function of time, this function being chosen in order to keep the maximum of the solution to the constant value 0. The main result of the note is…
We consider Hamilton Jacobi Bellman equations in an inifinite dimensional Hilbert space, with quadratic (respectively superquadratic) hamiltonian and with continuous (respectively lipschitz continuous) final conditions. This allows to study…
An abstract framework guaranteeing the local continuous differentiability of the value function associated with optimal stabilization problems subject to abstract semilinear parabolic equations subject to a norm constraint on the controls…
In this paper, we study the optimal stopping problem in the so-called exploratory framework, in which the agent takes actions randomly conditioning on current state and an entropy-regularized term is added to the reward functional. Such a…
We examine the problem of two-point boundary optimal control of nonlinear systems over finite-horizon time periods with unknown model dynamics by employing reinforcement learning. We use techniques from singular perturbation theory to…
The ergodic control problem for a non-degenerate controlled diffusion controlled through its drift is considered under a uniform stability condition that ensures the well-posedness of the associated Hamilton-Jacobi-Bellman (HJB) equation. A…
Policy iteration (PI) is a widely used algorithm for synthesizing optimal feedback control policies across many engineering and scientific applications. When PI is deployed on infinite-horizon, nonlinear, autonomous optimal-control…
The paper deals with a class of time-inconsistent control problems for McKean-Vlasov dynamics. By solving a backward time-inconsistent Hamilton-Jacobi-Bellman (HJB for short) equation coupled with a forward distribution-dependent stochastic…
An adaptive controller is proposed and analyzed for the class of infinite-horizon optimal control problems in positive linear systems presented in (Ohlin et al., 2024b). This controller is derived from the solution of a "data-driven…