Related papers: Neural Actor-Critic Methods for Hamilton-Jacobi-Be…
We study a structured bi-level optimization problem where the upper-level objective is a smooth function and the lower-level problem is policy optimization in a Markov decision process (MDP). The upper-level decision variable parameterizes…
This paper presents a mathematical formulation to perform temporal parallelisation of continuous-time optimal control problems, which can be solved via the Hamilton--Jacobi--Bellman (HJB) equation. We divide the time interval of the control…
We introduce a new and efficient numerical method for multicriterion optimal control and single criterion optimal control under integral constraints. The approach is based on extending the state space to include information on a "budget"…
A tensor decomposition approach for the solution of high-dimensional, fully nonlinear Hamilton-Jacobi-Bellman equations arising in optimal feedback control of nonlinear dynamics is presented. The method combines a tensor train approximation…
This work is devoted to the study of optimal control of stochastic functional differential equations (SFDEs) and its application to mathematical finance. By using the Dynkin formula and solution of the Dirichlet-Poisson problem, the…
We present a kernel-based linear matrix inequality (LMI) approach for the approximate solution of Hamilton--Jacobi--Bellman (HJB) equations arising in nonlinear optimal control. The method represents the gradient of the value function in a…
This paper deals with a family of stochastic control problems in Hilbert spaces which arises in typical applications (such as boundary control and control of delay equations with delay in the control) and for which is difficult to apply the…
Continuous-time reinforcement learning offers an appealing formalism for describing control problems in which the passage of time is not naturally divided into discrete increments. Here we consider the problem of predicting the distribution…
Multi-agent reinforcement learning has been successfully applied to a number of challenging problems. Despite these empirical successes, theoretical understanding of different algorithms is lacking, primarily due to the curse of…
This paper presents a novel method of global adaptive dynamic programming (ADP) for the adaptive optimal control of nonlinear polynomial systems. The strategy consists of relaxing the problem of solving the Hamilton-Jacobi-Bellman (HJB)…
In this paper, we focus on the stochastic representation of a system of coupled Hamilton-Jacobi-Bellman-Isaacs (HJB-Isaacs (HJBI), for short) equations which is in fact a system of coupled Isaacs' type integral-partial differential…
This study investigates a stochastic production planning problem with a running cost composed of quadratic production costs and inventory-dependent costs. The objective is to minimize the expected cost until production stops when inventory…
We establish, for the first time, explicit a priori and regularity estimates for solutions of the Dirichlet problem for Hamilton-Jacobi-Bellman operators from stochastic control, whose principal half-eigenvalues have opposite signs. In…
In this paper, we introduce Hamilton-Jacobi-Bellman (HJB) equations for Q-functions in continuous time optimal control problems with Lipschitz continuous controls. The standard Q-function used in reinforcement learning is shown to be the…
This paper presents an implicit solution formula for the Hamilton-Jacobi partial differential equation (HJ PDE). The formula is derived using the method of characteristics and is shown to coincide with the Hopf and Lax formulas in the case…
We study the properties of the value function associated with an optimal control problem with uncertainties, known as average or Riemann-Stieltjes problem. Uncertainties are assumed to belong to a compact metric probability space, and…
We address two major challenges in scientific machine learning (SciML): interpretability and computational efficiency. We increase the interpretability of certain learning processes by establishing a new theoretical connection between…
We introduce a novel extension to robust control theory that explicitly addresses uncertainty in the value function's gradient, a form of uncertainty endemic to applications like reinforcement learning where value functions are…
Deterministic policy gradient algorithms are foundational for actor-critic methods in controlling continuous systems, yet they often encounter inaccuracies due to their dependence on the derivative of the critic's value estimates with…
This paper establishes the existence and uniqueness of mild solutions to stationary Hamilton-Jacobi-Bellman (HJB) equations associated with infinite-horizon stochastic optimal control problems in separable Hilbert spaces. Our framework…