Related papers: Hamilton-Jacobi-Bellman Equations for Q-Learning i…
This paper, which is the natural continuation of a previous paper by the same authors, studies a class of optimal control problems with state constraints where the state equation is a differential equation with delays. This class includes…
This paper investigates the optimal control problems for the finite-horizon continuous-time Markov decision processes with delay-dependent control policies. We develop compactification methods in decision processes, and show that the…
We address two major challenges in scientific machine learning (SciML): interpretability and computational efficiency. We increase the interpretability of certain learning processes by establishing a new theoretical connection between…
In this paper we study a first extension of the theory of mild solutions for HJB equations in Hilbert spaces to the case when the domain is not the whole space. More precisely, we consider a half-space as domain, and a semilinear…
Optimal control and the associated second-order path-dependent Hamilton-Jacobi-Bellman (PHJB) equation are studied for unbounded functional stochastic evolution systems in Hilbert spaces. The notion of viscosity solution without…
Commonly in reinforcement learning (RL), rewards are discounted over time using an exponential function to model time preference, thereby bounding the expected long-term reward. In contrast, in economics and psychology, it has been shown…
In this paper, we study the optimal singular controls for stochastic recursive systems, in which the control has two components: the regular control, and the singular control. Under certain assumptions, we establish the dynamic programming…
This paper introduces the Hamilton-Jacobi-Bellman Proximal Policy Optimization (HJBPPO) algorithm into reinforcement learning. The Hamilton-Jacobi-Bellman (HJB) equation is used in control theory to evaluate the optimality of the value…
This note lays part of the theoretical ground for a definition of differential systems modeling reinforcement learning in continuous time non-Markovian rough environments. Specifically we focus on optimal relaxed control of rough equations…
In this paper, a stochastic optimal control problem is investigated in which the system is governed by a stochastic functional differential equation. In the framework of functional It\^o calculus, we build the dynamic programming principle…
Considering that the decision-making environment faced by reinforcement learning (RL) agents is full of Knightian uncertainty, this paper describes the exploratory state dynamics equation in Knightian uncertainty to study the…
Stochastic optimal control problems for Hamiltonian dynamics on graphs have wide-ranging applications in mechanics and quantum field theory, particularly in systems with graph-based structures. In this paper, we establish the existence and…
This work proposes a novel numerical scheme for solving the high-dimensional Hamilton-Jacobi-Bellman equation with a functional hierarchical tensor ansatz. We consider the setting of stochastic control, whereby one applies control to a…
In this paper, we study a time-inconsistent stochastic optimal control problem with a recursive cost functional by a multi-person hierarchical differential game approach. An equilibrium strategy of this problem is constructed and a…
Hamilton-Jacobi partial differential equations (HJ PDEs) have deep connections with a wide range of fields, including optimal control, differential games, and imaging sciences. By considering the time variable to be a higher dimensional…
This paper studies the stochastic optimal control of jump-diffusion processes and the associated fully nonlinear backward stochastic Hamilton--Jacobi--Bellman (BSHJB) equations. We establish the dynamic programming principle (DPP) via…
The Bellman equation and its continuous form, the Hamilton-Jacobi-Bellman equation, are ubiquitous in reinforcement learning and control theory. However, these equations become intractable for high-dimensional or nonlinear systems. This…
We study the well-posedness of Hamilton-Jacobi-Bellman equations on subsets of $\mathbb{R}^d$ in a context without boundary conditions. The Hamiltonian is given as the supremum over two parts: an internal Hamiltonian depending on an…
For an infinite-horizon control problem, the optimal control can be represented by the stable manifold of the characteristic Hamiltonian system of Hamilton-Jacobi-Bellman (HJB) equation in a semiglobal domain. In this paper, we first…
We study the problem of optimal portfolio selection under stochastic volatility within a continuous time reinforcement learning framework with portfolio constraints. Exploration is modeled through entropy-regularized relaxed controls, where…