Related papers: Continuous-Time Fitted Value Iteration for Robust …
We study the problem of learning the optimal control policy for fine-tuning a given diffusion process, using general value function approximation. We develop a new class of algorithms by solving a variational inequality problem based on the…
Feedback controllers for port-Hamiltonian systems reveal an intrinsic inverse optimality property since each passivating state feedback controller is optimal with respect to some specific performance index. Due to the nonlinear…
This paper is concerned with a stochastic recursive optimal control problem with time delay, where the controlled system is described by a stochastic differential delayed equation (SDDE) and the cost functional is formulated as the solution…
In this note, we study a class of indefinite stochastic McKean-Vlasov linear-quadratic (LQ in short) control problem under the control taking nonnegative values. In contrast to the conventional issue, both the classical dynamic programming…
This paper derives recursion equations for a robust smoothing problem for a class of nonlinear systems with uncertainties in modeling and exogenous noise sources. The systems considered operate in discrete-time and the uncertainties are…
Despite impressive results, reinforcement learning (RL) suffers from slow convergence and requires a large variety of tuning strategies. In this paper, we investigate the ability of RL algorithms on simple continuous control tasks. We show…
We study a continuous-time portfolio choice problem for an investor whose state-dependent preferences are determined by an exogenous factor that evolves as an It\^o diffusion process. Since risk attitudes at the end of the investment…
For a general entropy-regularized time-inconsistent stochastic control problem, we propose a policy iteration algorithm (PIA) and establish its convergence to an equilibrium policy with an exponential convergence rate. The design of the PIA…
In this paper, we study a novel episodic risk-sensitive Reinforcement Learning (RL) problem, named Iterated CVaR RL, which aims to maximize the tail of the reward-to-go at each step, and focuses on tightly controlling the risk of getting…
Deterministic optimal impulse control problem with terminal state constraint is considered. Due to the appearance of the terminal state constraint, the value function might be discontinuous in general. The main contribution of this paper is…
This paper proposes penalty schemes for a class of weakly coupled systems of Hamilton-Jacobi-Bellman quasi-variational inequalities (HJBQVIs) arising from stochastic hybrid control problems of regime-switching models with both continuous…
Reinforcement learning based adaptive/approximate dynamic programming (ADP) is a powerful technique to determine an approximate optimal controller for a dynamical system. These methods bypass the need to analytically solve the nonlinear…
Approximate dynamic programming algorithms, such as approximate value iteration, have been successfully applied to many complex reinforcement learning tasks, and a better approximate dynamic programming algorithm is expected to further…
The optimal \(H_{\infty}\) control problem over an infinite time horizon, which incorporates a performance function with a discount factor \(e^{-\alpha t}\) (\(\alpha > 0\)), is important in various fields. Solving this optimal…
We propose a class of numerical schemes for nonlocal HJB variational inequalities (HJBVIs) with monotone drivers. The solution and free boundary of the HJBVI are constructed from a sequence of penalized equations, for which a continuous…
The aim of this work is to develop a deep learning method for solving high-dimensional stochastic control problems based on the Hamilton--Jacobi--Bellman (HJB) equation and physics-informed learning. Our approach is to parameterize the…
In this paper, we study a time-inconsistent stochastic optimal control problem with a recursive cost functional by a multi-person hierarchical differential game approach. An equilibrium strategy of this problem is constructed and a…
In a recent work, we proposed Reliable Policy Iteration (RPI), that restores policy iteration's monotonicity-of-value-estimates property to the function approximation setting. Here, we assess the robustness of RPI's empirical performance on…
In the present paper, we study a two-player zero-sum deterministic differential game with both players adopting impulse controls, in infinite time horizon, under rather weak assumptions on the cost functions. We prove by means of the…
We develop a computationally efficient learning-based forward-backward stochastic differential equations (FBSDE) controller for both continuous and hybrid dynamical (HD) systems subject to stochastic noise and state constraints. Solutions…