Related papers: The Ces\`aro Value Iteration
For a general entropy-regularized stochastic control problem on an infinite horizon, we prove that a policy iteration algorithm (PIA) converges to an optimal relaxed control. Contrary to the standard stochastic control literature, classical…
An iterative learning algorithm is presented for continuous-time linear-quadratic optimal control problems where the system is externally symmetric with unknown dynamics. Both finite-horizon and infinite-horizon problems are considered. It…
In this paper, infinite horizon stochastic difference equations and backward stochastic difference equations with fractional noises are studied. The main difficulty comes from fractional noises on infinite horizon. Motivated by…
This paper is concerned with a time-inconsistent stochastic optimal control problem in an infinite time horizon with a non-degenerate diffusion in the state equation. A major assumption is that people become rational after a large time.…
This paper studies value iteration for infinite horizon contracting Markov decision processes under convexity assumptions and when the state space is uncountable. The original value iteration is replaced with a more tractable form and the…
In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed…
This paper is devoted to a study of infinite horizon optimal control problems with time discounting and time averaging criteria in discrete time. We establish that these problems are related to certain infinite-dimensional linear…
In this paper, a finite-horizon optimal control problem involving a dynamical system described by a linear Caputo fractional differential equation and a quadratic cost functional is considered. An explicit formula for the value functional…
This paper is devoted to the analysis of a finite horizon discrete-time stochastic optimal control problem, in presence of constraints. We study the regularity of the value function which comes from the dynamic programming algorithm. We…
We study the problem of computing optimal correlated equilibria (CEs) in infinite-horizon multi-player stochastic games, where correlation signals are provided over time. In this setting, optimal CEs require history-dependent policies; this…
The Inverse Optimal Control (IOC) problem is a structured system identification problem that aims to identify the underlying objective function based on observed optimal trajectories. This provides a data-driven way to model experts'…
In this paper, we propose a new policy iteration algorithm to compute the value function and the optimal controls of continuous time stochastic control problems. The algorithm relies on successive approximations using linear-quadratic…
We propose a machine learning algorithm for solving finite-horizon stochastic control problems based on a deep neural network representation of the optimal policy functions. The algorithm has three features: (1) It can solve…
We present a theory of optimal control for McKean-Vlasov stochastic differential equations with infinite time horizon and discounted gain functional. We first establish the well-posedness of the state equation and of the associated control…
This paper presents an iterative learning control (ILC) scheme for continuously operated repetitive systems for which no initial condition reset exists. To accomplish this, we develop a lifted system representation that accounts for the…
This paper is concerned with a stochastic linear quadratic (LQ, for short) control problem with a recursive cost functional in an infinite horizon. A main difficult is well-posedness of the BSDE in $L^1$ and in infinite horizon. A notion of…
The present paper is devoted to the study of the asymptotic behavior of the value functions of both finite and infinite horizon stochastic control problems and to the investigation of their relation with suitable stochastic ergodic control…
We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework…
We consider impulse control problems in finite horizon for diffusions with decision lag and execution delay. The new feature is that our general framework deals with the important case when several consecutive orders may be decided before…
We design receding horizon control strategies for stochastic discrete-time linear systems with additive (possibly) unbounded disturbances, while obeying hard bounds on the control inputs. We pose the problem of selecting an appropriate…