Related papers: Pathwise Relaxed Optimal Control of Rough Differen…
In recent times, reinforcement learning has produced baffling results when it comes to performing control tasks with highly non-linear systems. The impressive results always outweigh the potential vulnerabilities or uncertainties associated…
In the present work we employ, for the first time, backward stochastic differential equations (BSDEs) to study the optimal control of semi-Markov processes on finite horizon, with general state and action spaces. More precisely, we prove…
The paper concerns the infinite dimensional Hamilton-Jacobi-Bellman equation related to optimal control problem regulated by a transport equation with boundary control. A suitable viscosity solution approach is needed in view of the…
In this work we investigate regularity properties of a large class of Hamilton-Jacobi-Bellman (HJB) equations with or without obstacles, which can be stochastically interpreted in form of a stochastic control system which nonlinear cost…
This paper presents a physics-informed machine learning approach for synthesizing optimal feedback control policy for infinite-horizon optimal control problems by solving the Hamilton-Jacobi-Bellman (HJB) partial differential equation(PDE).…
Merely pursuing performance may adversely affect the safety, while a conservative policy for safe exploration will degrade the performance. How to balance the safety and performance in learning-based control problems is an interesting yet…
The entropy regularization is inspired by information entropy from machine learning and the ideas of exploration and exploitation in reinforcement learning, which appears in the control problem to design an approximating algorithm for the…
This paper introduces a new type of second order stochastic backward Hamilton-Jacobi-Bellman (HJB) equations for optimal stochastic control problems with a currently observable but non-predicable parameter process, in addition to the…
This work uses the entropy-regularised relaxed stochastic control perspective as a principled framework for designing reinforcement learning (RL) algorithms. Herein agent interacts with the environment by generating noisy controls…
This paper studies the optimal dividend problem with a bounded payout rate in a partially observed regime-switching diffusion model, where, in practice, the market regime is unobserved and key model parameters are unknown. To address this…
In this paper we investigate a kind of optimal control problem of coupled forward-backward stochastic system with jumps whose cost functional is defined through a coupled forward-backward stochastic differential equation with Brownian…
This paper presents a new method for synthesizing stochastic control Lyapunov functions for a class of nonlinear stochastic control systems. The technique relies on a transformation of the classical nonlinear Hamilton-Jacobi-Bellman partial…
This work addresses an optimal control problem constrained by a degenerate kinetic equation of parabolic-hyperbolic type. Using a hypocoercivity framework we establish the well-posedness of the problem and demonstrate that the optimal…
Inverse reinforcement learning aims to infer the reward function that explains expert behavior observed through trajectories of state--action pairs. A long-standing difficulty in classical IRL is the non-uniqueness of the recovered reward:…
We consider a stochastic control problem where the set of controls is not necessarily convex and the system is governed by a nonlinear backward stochastic differential equation. We establish necessary as well as sufficient conditions of…
We study the properties of the value function associated with an optimal control problem with uncertainties, known as average or Riemann-Stieltjes problem. Uncertainties are assumed to belong to a compact metric probability space, and…
We deal with the convergence of the value function of an approximate control problem with uncertain dynamics to the value function of a nonlinear optimal control problem. The assumptions on the dynamics and the costs are rather general and…
We show that the value function of a stochastic control problem is the unique solution of the associated Hamilton-Jacobi-Bellman (HJB) equation, completely avoiding the proof of the so-called dynamic programming principle (DPP). Using…
This paper presents a novel model-reference reinforcement learning control method for uncertain autonomous surface vehicles. The proposed control combines a conventional control method with deep reinforcement learning. With the conventional…
In this paper we prove that there exists a smooth classical solution to the HJB equation for a large class of constrained problems with utility functions that are not necessarily differentiable or strictly concave. The value function is…