Related papers: A Temporal Difference Method for Stochastic Contin…
For continuous action spaces, actor-critic methods are widely used in online reinforcement learning (RL). However, unlike RL algorithms for discrete actions, which generally model the optimal value function using the Bellman optimality…
Learning-based control methods typically assume stationary system dynamics, an assumption often violated in real-world systems due to drift, wear, or changing operating conditions. We study reinforcement learning for control under…
This paper investigates the optimal control problems for the finite-horizon continuous-time Markov decision processes with delay-dependent control policies. We develop compactification methods in decision processes, and show that the…
We present an accelerated algorithm for the solution of static Hamilton-Jacobi-Bellman equations related to optimal control problems. Our scheme is based on a classic policy iteration procedure, which is known to have superlinear…
In this paper, we attempt to introduce the Bellman principle for a discrete time multi-period mean-variance model. Based on this new take on the Bellman principle, we obtain a dynamic time-consistent optimal strategy and related efficient…
A learning technique for finite horizon optimal control problems and its approximation based on polynomials is analyzed. It allows to circumvent, in part, the curse dimensionality which is involved when the feedback law is constructed by…
In this work we investigate regularity properties of a large class of Hamilton-Jacobi-Bellman (HJB) equations with or without obstacles, which can be stochastically interpreted in form of a stochastic control system which nonlinear cost…
This paper proposes a new reinforcement learning with hyperbolic discounting. Combining a new temporal difference error with the hyperbolic discounting in recursive manner and reward-punishment framework, a new scheme to learn the optimal…
In this paper, which is a continuation of the previously published discrete time paper we develop a theory for continuous time stochastic control problems which, in various ways, are time inconsistent in the sense that they do not admit a…
We apply the Stochastic Perron method, created by Bayraktar and S\^irbu, to a stochastic exit time control problem. Our main assumption is the validity of the Strong Comparison Result for the related Hamilton-Jacobi-Bellman (HJB) equation.…
We develop a computationally efficient learning-based forward-backward stochastic differential equations (FBSDE) controller for both continuous and hybrid dynamical (HD) systems subject to stochastic noise and state constraints. Solutions…
The State-Dependent Riccati Equation (SDRE) approach is extensively utilized in nonlinear optimal control as a reliable framework for designing robust feedback control strategies. This work provides an analysis of the SDRE approach,…
We study linear-quadratic stochastic optimal control problems with bilinear state dependence for which the underlying stochastic differential equation (SDE) consists of slow and fast degrees of freedom. We show that, in the same way in…
This paper considers consumption and portfolio optimization problems with recursive preferences in both infinite and finite time regions. Specially, the financial market consists of a risk-free asset and a risky asset that follows a general…
Hamilton-Jacobi reachability (HJR) is an exciting framework used for control of safety-critical systems with nonlinear and possibly uncertain dynamics. However, HJR suffers from the curse of dimensionality, with computation times growing…
We derive an equation for temporal difference learning from statistical principles. Specifically, we start with the variational principle and then bootstrap to produce an updating rule for discounted state value estimates. The resulting…
We study the convergence of Markov Decision Processes made of a large number of objects to optimization problems on ordinary differential equations (ODE). We show that the optimal reward of such a Markov Decision Process, satisfying a…
A new framework for formulating reachability problems with competing inputs, nonlinear dynamics and state constraints as optimal control problems is developed. Such reach-avoid problems arise in, among others, the study of safety problems…
Feedback controllers for port-Hamiltonian systems reveal an intrinsic inverse optimality property since each passivating state feedback controller is optimal with respect to some specific performance index. Due to the nonlinear…
In this paper we study a class of stochastic control problems in which the control of the jump size is essential. Such a model is a generalized version for various applied problems ranging from optimal reinsurance selections for general…