Related papers: Long run stochastic control problems with general …
This paper is devoted to studying the average optimality in continuous-time Markov decision processes with fairly general state and action spaces. The criterion to be maximized is expected average rewards. The transition rates of underlying…
We address the variational formulation of the risk-sensitive reward problem for non-degenerate diffusions on $\mathbb{R}^d$ controlled through the drift. We establish a variational formula on the whole space and also show that the…
This paper firstly presents the necessary and sufficient conditions for a kind of discrete-time robust stochastic optimal control problem with convex control domains. As it is an "inf sup problem", the classical variational method is…
A G/M/N queue is considered in the moderate deviation heavy traffic regime. The rate function for the customers-in-system process is obtained for the single class model. A risk-sensitive type control problem is considered for multi-class…
Reflected diffusions naturally arise in many problems from applications ranging from economics and mathematical biology to queueing theory. In this paper we consider a class of infinite time-horizon singular stochastic control problems for…
Traditional reinforcement learning (RL) aims to maximize the expected total reward, while the risk of uncertain outcomes needs to be controlled to ensure reliable performance in a risk-averse setting. In this paper, we consider the problem…
This paper is concerned with the solution of the optimal stopping problem associated to the valuation of Perpetual American options driven by continuous time Markov chains. We introduce a new dynamic approach for the numerical pricing of…
We study Markov decision processes with Polish state and action spaces. The action space is state dependent and is not necessarily compact. We first establish the existence of an optimal ergodic occupation measure using only a near-monotone…
We address the problem of policy evaluation in discounted Markov decision processes, and provide instance-dependent guarantees on the $\ell_\infty$-error under a generative model. We establish both asymptotic and non-asymptotic versions of…
Motivated by robotic surveillance applications, this paper studies the novel problem of maximizing the return time entropy of a Markov chain, subject to a graph topology with travel times and stationary distribution. The return time entropy…
This paper addresses objectives tailored to the risk-averse optimization of accumulated rewards in Markov decision processes (MDPs). The studied objectives require maximizing the expected value of the accumulated rewards minus a penalty…
This paper considers receding horizon control of finite deterministic systems, which must satisfy a high level, rich specification expressed as a linear temporal logic formula. Under the assumption that time-varying rewards are associated…
We consider the challenge of finding a deterministic policy for a Markov decision process that uniformly (in all states) maximizes one reward subject to a probabilistic constraint over a different reward. Existing solutions do not fully…
This paper is devoted to studying constrained continuous-time Markov decision processes (MDPs) in the class of randomized policies depending on state histories. The transition rates may be unbounded, the reward and costs are admitted to be…
In the classical static optimal reinsurance problem, the cost of capital for the insurer's risk exposure determined by a monetary risk measure is minimized over the class of reinsurance treaties represented by increasing Lipschitz retained…
A Markov decision process can be parameterized by a transition kernel and a reward function. Both play essential roles in the study of reinforcement learning as evidenced by their presence in the Bellman equations. In our inquiry of various…
We aim at characterizing the asymptotic behavior of value functions in the control of piece-wise deterministic Markov processes (PDMP) of switch type under nonexpansive assumptions. For a particular class of processes inspired by temperate…
We investigate the well-posedness of a general class of singular stochastic control problems in which controls are processes of finite variation. We develop an abstract framework, which we then apply to storage management and portfolio…
Consider the following multi-phase project management problem. Each project is divided into several phases. All projects enter the next phase at the same point chosen by the decision maker based on observations up to that point. Within each…
We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove…