Related papers: Undiscounted optimal stopping with unbounded rewar…
We study one-sided and $\alpha$-correct sequential hypothesis testing for data generated by an ergodic Markov chain. The null hypothesis is that the unknown transition matrix belongs to a prescribed set $P$ of stochastic matrices, and the…
We present a method to solve optimal stopping problems in infinite horizon for a L\'evy process when the reward function can be non-monotone. To solve the problem we introduce two new objects. Firstly, we define a random variable $\eta(x)$…
The problem of optimal stopping with finite horizon in discrete time is considered in view of maximizing the expected gain. The algorithm proposed in this paper is completely nonparametric in the sense that it uses observed data from the…
We study the properties of the free boundaries and the corresponding hitting times in the context of optimal stopping in discrete time. We first prove the continuity of the map from the boundaries to the expected value of the corresponding…
In this paper we consider discrete and continuous time risk sensitive optimal stopping problem. Using suitable properties of the underlying Feller-Markov process we prove continuity of the optimal stopping value function and provide formula…
In constrained Markov decision processes (CMDPs) with adversarial rewards and constraints, a well-known impossibility result prevents any algorithm from attaining both sublinear regret and sublinear constraint violation, when competing…
We study the optimal multiple stopping time problem defined for each stopping time $S$ by $v(S)=\operatorname {ess}\sup_{\tau_1,...,\tau_d\geq S}E[\psi(\tau_1,...,\tau_d)|\mathcal{F}_S]$. The key point is the construction of a new reward…
The paper addresses the problem of computing maximal conditional expected accumulated rewards until reaching a target state (briefly called maximal conditional expectations) in finite-state Markov decision processes where the condition is…
For a discrete time Markov chain and in line with Strotz' consistent planning we develop a framework for problems of optimal stopping that are time-inconsistent due to the consideration of a non-linear function of an expected reward. We…
We consider optimal stopping problems, in which a sequence of independent random variables is drawn from a known continuous density. The objective of such problems is to find a procedure which maximizes the expected reward; this is often…
A general result on the method of randomized stopping is proved. It is applied to optimal stopping of controlled diffusion processes with unbounded coefficients to reduce it to an optimal control problem without stopping. This is motivated…
Markov chains are the de facto finite-state model for stochastic dynamical systems, and Markov decision processes (MDPs) extend Markov chains by incorporating non-deterministic behaviors. Given an MDP and rewards on states, a classical…
This paper is devoted to studying constrained continuous-time Markov decision processes (MDPs) in the class of randomized policies depending on state histories. The transition rates may be unbounded, the reward and costs are admitted to be…
We consider mean-field control problems in discrete time with discounted reward, infinite time horizon and compact state and action space. The existence of optimal policies is shown and the limiting mean-field problem is derived when the…
We present a method to find an optimal policy with respect to a reward function for a discounted Markov decision process under general linear temporal logic (LTL) specifications. Previous work has either focused on maximizing a cumulative…
Standard Markovian optimal stopping problems are consistent in the sense that the first entrance time into the stopping set is optimal for each initial state of the process. Clearly, the usual concept of optimality cannot in a…
Markov reward processes (MRPs) are used to model stochastic phenomena arising in operations research, control engineering, robotics, and artificial intelligence, as well as communication and transportation networks. In many of these cases,…
In this paper, we study the optimal stopping problem in the so-called exploratory framework, in which the agent takes actions randomly conditioning on current state and an entropy-regularized term is added to the reward functional. Such a…
Markov decision processes (MDPs) with rewards are a widespread and well-studied model for systems that make both probabilistic and nondeterministic choices. A fundamental result about MDPs is that their minimal and maximal expected rewards…
We present the first finite time global convergence analysis of policy gradient in the context of infinite horizon average reward Markov decision processes (MDPs). Specifically, we focus on ergodic tabular MDPs with finite state and action…