Related papers: Is Bellman Equation Enough for Learning Control?
We determine the optimal robust investment strategy of an individual who targets at a given rate of consumption and seeks to minimize the probability of lifetime ruin when she does not have perfect confidence in the drift of the risky…
We introduce a novel extension to robust control theory that explicitly addresses uncertainty in the value function's gradient, a form of uncertainty endemic to applications like reinforcement learning where value functions are…
We study a discounted singular stochastic control problem driven by a general L\'evy process, where the objective is to minimize a cost functional composed of a running cost and a control cost that depends on the current state of the…
We solve in mild sense Hamilton Jacobi Bellman equations, both in an infinite dimensional Hilbert space and in a Banach space, with lipschitz Hamiltonian and lipschitz continuous final condition, and asking only a weak regularizing property…
This paper introduces a new type of second order stochastic backward Hamilton-Jacobi-Bellman (HJB) equations for optimal stochastic control problems with a currently observable but non-predicable parameter process, in addition to the…
In recent work it is shown that Q-learning with linear function approximation is stable, in the sense of bounded parameter estimates, under the $(\varepsilon,\kappa)$-tamed Gibbs policy; $\kappa$ is inverse temperature, and $\varepsilon>0$…
We introduce a new numerical method to approximate the solution of a finite horizon deterministic optimal control problem. We exploit two Hamilton-Jacobi-Bellman PDE, arising by considering the dynamics in forward and backward time. This…
We develop the dynamic programming approach for a family of infinite horizon boundary control problems with linear state equation and convex cost. We prove that the value function of the problem is the unique regular solution of the…
We consider a stochastic control problem with the assumption that the system is controlled until the state process breaks the fixed barrier. Assuming some general conditions, it is proved that the resulting Hamilton Jacobi Bellman equations…
We devise a control-theoretic reinforcement learning approach to support direct learning of the optimal policy. We establish various theoretical properties of our approach, such as convergence and optimality of our analog of the Bellman…
We study the structure of a simple dynamic optimization problem consisting of one state and one control variable, from a physicist's point of view. By using an analogy to a physical model, we study this system in the classical and quantum…
This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type of exploratory formulation under entropy regularization where the agent randomizes both the timing…
Optimal control problem is typically solved by first finding the value function through Hamilton-Jacobi equation (HJE) and then taking the minimizer of the Hamiltonian to obtain the control. In this work, instead of focusing on the value…
In reinforcement learning, temporal difference-based algorithms can be sample-inefficient: for instance, with sparse rewards, no learning occurs until a reward is observed. This can be remedied by learning richer objects, such as a model of…
The purpose of this work is to introduce a notion of weak solution to the master equation of a potential mean field game and to prove that existence and uniqueness hold under quite general assumptions. Remarkably, this is achieved without…
In this paper, we present a novel algorithm named synchronous integral Q-learning, which is based on synchronous policy iteration, to solve the continuous-time infinite horizon optimal control problems of input-affine system dynamics. The…
Some approaches to solving challenging dynamic programming problems, such as Q-learning, begin by transforming the Bellman equation into an alternative functional equation, in order to open up a new line of attack. Our paper studies this…
Supervised machine learning is powerful. In recent years, it has enabled massive breakthroughs in computer vision and natural language processing. But leveraging these advances for optimal control has proved difficult. Data is a key…
In this work, we consider the local Cahn-Hilliard-Navier-Stokes equation with regular potential in two dimensional bounded domain. We formulate distributed optimal control problem as the minimization of a suitable cost functional subject to…
In this paper, we consider the stochastic optimal control problem for jump diffusion systems with state constraints. In general, the value function of such problems is a discontinuous viscosity solution of the Hamilton-Jacobi-Bellman (HJB)…