Related papers: Mean field for Markov Decision Processes: from Dis…
This paper, which is the natural continuation of a previous paper by the same authors, studies a class of optimal control problems with state constraints where the state equation is a differential equation with delays. This class includes…
We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains…
A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…
We study a class of optimal control problems with state constraints where the state equation is a differential equation with delays. This class includes some problems arising in economics, in particular the so-called models with time to…
Stochastic domains often involve risk-averse decision makers. While recent work has focused on how to model risk in Markov decision processes using risk measures, it has not addressed the problem of solving large risk-averse formulations.…
We introduce universal, easy-to-reproduce generative models for the QUBO instances to differentiate the performance of the hardware/solvers effectively. Our benchmark process extends the well-known Hebb's rule of associative memory with the…
Calculating optimal policies is known to be computationally difficult for Markov decision processes (MDPs) with Borel state and action spaces. This paper studies finite-state approximations of discrete time Markov decision processes with…
We study optimality for the safety-constrained Markov decision process which is the underlying framework for safe reinforcement learning. Specifically, we consider a constrained Markov decision process (with finite states and finite…
In this paper, we study closed-loop equilibrium strategies for mean-variance portfolio selection problem in a hidden Markov model with dynamic attention behavior. In addition to the investment strategy, the investor's attention to news is…
The long-run average payoff per transition (mean payoff) is the main tool for specifying the performance and dependability properties of discrete systems. The problem of constructing a controller (strategy) simultaneously optimizing several…
We introduce a novel approach to hierarchical reinforcement learning for Linearly-solvable Markov Decision Processes (LMDPs) in the infinite-horizon average-reward setting. Unlike previous work, our approach allows learning low-level and…
The centralized training for decentralized execution paradigm emerged as the state-of-the-art approach to $\epsilon$-optimally solving decentralized partially observable Markov decision processes. However, scalability remains a significant…
We consider a time-consistent mean-variance portfolio selection problem of an insurer and allow for the incorporation of basis (mortality) risk. The optimal solution is identified with a Nash subgame perfect equilibrium. We characterize an…
There are two primary approaches to solving Markov decision problems (MDPs): dynamic programming based on the Bellman equation and linear programming (LP). Dynamic programming methods are the most widely used and form the foundation of both…
Markov Decision Problems (MDPs) provide a foundational framework for modelling sequential decision-making across diverse domains, guided by optimality criteria such as discounted and average rewards. However, these criteria have inherent…
We consider a portfolio optimization problem in a defaultable market with finitely-many economical regimes, where the investor can dynamically allocate her wealth among a defaultable bond, a stock, and a money market account. The market…
We show that necessary and sufficient conditions of optimality in periodic optimization problems can be stated in terms of a solution of the corresponding HJB inequality, the latter being equivalent to a max-min type variational problem…
This article studies a portfolio optimization problem, where the market consisting of several stocks is modeled by a multi-dimensional jump-diffusion process with age-dependent semi-Markov modulated coefficients. We study risk sensitive…
In this article, we provide a numerical method based on fitted finite volume method to approximate the Hamilton-Jacobi-Bellman (HJB) equation coming from stochastic optimal control problems. The computational challenge is due to the nature…
Multi-agent reinforcement learning methods have shown remarkable potential in solving complex multi-agent problems but mostly lack theoretical guarantees. Recently, mean field control and mean field games have been established as a…