Related papers: Is Bellman Equation Enough for Learning Control?
It is well known that the extension of Watkins' algorithm to general function approximation settings is challenging: does the projected Bellman equation have a solution? If so, is the solution useful in the sense of generating a good…
An abstract framework guaranteeing the continuous differentiability of local value functions on $H^1(\Omega)$ associated with optimal stabilization problems subject to abstract semilinear parabolic equations in the presence of norm…
The objective of designing a control system is to steer a dynamical system with a control signal, guiding it to exhibit the desired behavior. The Hamilton-Jacobi-Bellman (HJB) partial differential equation offers a framework for optimal…
Learning optimal feedback control laws capable of executing optimal trajectories is essential for many robotic applications. Such policies can be learned using reinforcement learning or planned using optimal control. While reinforcement…
We study a time-optimal control problem of a two-peakon collision. First, we state the controllability. Next, we find the time-optimal strategy. This is done via the HamiltonJacobi-Bellman equation and the dynamic programming method. We…
Model-free algorithms for reinforcement learning typically require a condition called Bellman completeness in order to successfully operate off-policy with function approximation, unless additional conditions are met. However, Bellman…
This manuscript studies the Minkowski-Bellman equation, which is the Bellman equation arising from finite or infinite horizon optimal control of unconstrained linear discrete time systems with stage and terminal cost functions specified as…
We develop a discrete analogue of Hamilton-Jacobi theory in the framework of discrete Hamiltonian mechanics. The resulting discrete Hamilton-Jacobi equation is discrete only in time. We describe a discrete analogue of Jacobi's solution and…
Choosing how much noise to add in Langevin dynamics is essential for making these algorithms effective in challenging optimization problems. One promising approach is to determine this noise by solving Hamilton-Jacobi-Bellman (HJB)…
Reliable long-horizon value prediction is difficult in offline reinforcement learning because fitted value methods combine bootstrapping, function approximation, and distribution shift, while standard guarantees often require Bellman…
We study a multiscale stochastic optimal control problem subject to state constraints on the slow variable. To address this class of problems, we develop a rigorous theoretical framework based on singular perturbation analysis, tailored to…
Recent results in the study of the Hamilton Jacobi Bellman (HJB) equation have led to the discovery of a formulation of the value function as a linear Partial Differential Equation (PDE) for stochastic nonlinear systems with a mild…
Optimal control and the associated second-order path-dependent Hamilton-Jacobi-Bellman (PHJB) equation are studied for unbounded functional stochastic evolution systems in Hilbert spaces. The notion of viscosity solution without…
We present a new formulation for the computation of solutions of a class of Hamilton Jacobi Bellman (HJB) equations on closed smooth surfaces of co-dimension one. For the class of equations considered in this paper, the viscosity solution…
This paper studies the time-inconsistent MV optimal stopping problem via a game-theoretic approach to find equilibrium strategies. To overcome the mathematical intractability of direct equilibrium analysis, we propose a vanishing…
We consider an extension of the well-known Hamilton-Jacobi-Bellman (HJB) equation for fractional order dynamical systems in which a generalized performance index is considered for the related optimal control problem. Owing to the…
In this paper we study the existence of sufficiently regular representations of Hamilton-Jacobi equations in the optimal control theory with unbounded control set. We use a new method to construct representations for a wide class of…
Hard constraints in reinforcement learning (RL) often degrade policy performance. Lagrangian methods offer a way to blend objectives with constraints, but require intricate reward engineering and parameter tuning. In this work, we extend…
This paper deals with junction conditions for Hamilton-Jacobi-Bellman (HJB) equations for finite horizon control problems on multi-domains. We consider two different cases where the final cost is continuous or lower semi-continuous. In the…
We formulate a path-dependent stochastic optimal control problem under general conditions, for which weprove rigorously the dynamic programming principle and that the value function is the unique Crandall-Lions viscosity solution of the…