Related papers: Mean field for Markov Decision Processes: from Dis…
In this paper we propose a new way of proving the value of a firm that is currently producing a certain product and faces the option to exit the market. The problem of optimal exiting is an optimal stopping problem, that can be solved using…
In this paper, we present a mean field game to model the production behaviors of a very large number of producers, whose carbon emissions are regulated by government. Especially, an emission permits trading scheme is considered in our…
In this paper we study a continuous time equilibrium model of limit order book (LOB) in which the liquidity dynamics follows a non-local, reflected mean-field stochastic differential equation (SDE) with evolving intensity. Generalizing the…
We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove…
We model the stock price dynamics through a semi-Markov process obtained using a Poisson random measure. We establish the existence and uniqueness of the classical solution of a non-homogeneous terminal value problem and we show that the…
We propose a price impact model where changes in prices are purely driven by the order flow in the market. The stochastic price impact of market orders and the arrival rates of limit and market orders are functions of the market liquidity…
We study multi-objective reinforcement learning with nonlinear preferences over trajectories. That is, we maximize the expected value of a nonlinear function over accumulated rewards (expected scalarized return or ESR) in a multi-objective…
To sidestep the curse of dimensionality when computing solutions to Hamilton-Jacobi-Bellman partial differential equations (HJB PDE), we propose an algorithm that leverages a neural network to approximate the value function. We show that…
Robust Markov decision processes (MDPs) address the challenge of model uncertainty by optimizing the worst-case performance over an uncertainty set of MDPs. In this paper, we focus on the robust average-reward MDPs under the model-free…
We study a class of backward stochastic differential equations (BSDEs) driven by a random measure or, equivalently, by a marked point process. Under appropriate assumptions we prove well-posedness and continuous dependence of the solution…
We study value-iteration (VI) algorithms for solving general (a.k.a. multichain) Markov decision processes (MDPs) under the average-reward criterion, a fundamental but theoretically challenging setting. Beyond the difficulties inherent to…
We present a simple and easy to implement method for the numerical solution of a rather general class of Hamilton-Jacobi-Bellman (HJB) equations. In many cases, the considered problems have only a viscosity solution, to which, fortunately,…
In this paper we study backward stochastic differential equations (BSDEs) driven by the compensated random measure associated to a given pure jump Markov process X on a general state space K. We apply these results to prove well-posedness…
Discrete time stochastic optimal control problems and Markov decision processes (MDPs), respectively, serve as fundamental models for problems that involve sequential decision making under uncertainty and as such constitute the theoretical…
Mean field game facilitates analyzing multi-armed bandit (MAB) for a large number of agents by approximating their interactions with an average effect. Existing mean field models for multi-agent MAB mostly assume a binary reward function,…
In reinforcement learning (RL), aligning agent behavior with specific objectives typically requires careful design of the reward function, which can be challenging when the desired objectives are complex. In this work, we propose an…
A class of stochastic optimal control problems involving optimal stopping is considered. Methods of Krylov are adapted to investigate the numerical solutions of the corresponding normalized Bellman equations and to estimate the rate of…
Motion planning under uncertainty for an autonomous system can be formulated as a Markov Decision Process with a continuous state space. In this paper, we propose a novel solution to this decision-theoretic planning problem that directly…
The Ordered Upwind Method (OUM) is used to approximate the viscosity solution of the static Hamilton-Jacobi-Bellman (HJB) with direction-dependent weights on unstructured meshes. The method has been previously shown to provide a solution…
Controlling systems of ordinary differential equations (ODEs) is ubiquitous in science and engineering. For finding an optimal feedback controller, the value function and associated fundamental equations such as the Bellman equation and the…