Related papers: Undiscounted optimal stopping with unbounded rewar…
We extend the classical setting of an optimal stopping problem under full information to include for problems with an unknown state. The framework allows the unknown state to influence (i) the drift of the underlying process, (ii) the…
We conduct an investigation of the differentiability and continuity of reward functionals associated to Markovian randomized stopping times. Our focus is mostly on the differentiability, which is a crucial ingredient for a common approach…
In this Note we study optimal stopping problems for strong Markov processes and affine functions. We give a justification of the Snell envelope form using standard results of optimal stopping. We also justify the convexity of the value…
We analyze an optimal stopping problem with random maturity under a nonlinear expectation with respect to a weakly compact set of mutually singular probabilities $\mathcal{P}$. The maturity is specified as the hitting time to level $0$ of…
This paper considers the infinite horizon optimal control problem for nonlinear systems. Under the condition of nonlinear controllability of the system to any terminal set containing the origin and forward invariance of the terminal set, we…
This paper deals with the unconstrained and constrained cases for continuous-time Markov decision processes under the finite-horizon expected total cost criterion. The state space is denumerable and the transition and cost rates are allowed…
We consider the optimal stopping problem for a Gauss-Markov process conditioned to adopt a prescribed terminal distribution. By applying a time-space transformation, we show it is equivalent to stopping a Brownian bridge pinned at a random…
This paper studies the expected value of multiplicative rewards, where rewards obtained in each step are multiplied (instead of the usual addition), in Markov chains (MCs) and Markov decision processes (MDPs). One of the key differences to…
A class of stochastic optimal control problems involving optimal stopping is considered. Methods of Krylov are adapted to investigate the numerical solutions of the corresponding normalized Bellman equations and to estimate the rate of…
We consider the Gittins index for a normal distribution with unknown mean $\theta$ and known variance where $\theta$ has a normal prior. In addition to presenting some monotonicity properties of the Gittins index, we derive an approximation…
Our purpose is to study a particular class of optimal stopping problems for Markov processes. We justify the value function convexity and we deduce that there exists a boundary function such that the smallest optimal stopping time is the…
The best arm identification problem requires identifying the best alternative (i.e., arm) in active experimentation using the smallest number of experiments (i.e., arm pulls), which is crucial for cost-efficient and timely decision-making…
In this paper, we investigate the concentration properties of cumulative reward in Markov Decision Processes (MDPs), focusing on both asymptotic and non-asymptotic settings. We introduce a unified approach to characterize reward…
We consider the constrained optimal control problem for the gradual-impulsive CTMDP model with the performance criteria being the expected total undiscounted costs (from the running cost and the cost from each time an impulse being…
A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or…
We propose a numerical method to approximate the value function for the optimal stopping problem of a piecewise deterministic Markov process (PDMP). Our approach is based on quantization of the post jump location---inter-arrival time Markov…
A recent work by Schlisselberg et al. (2024) studies a delay-as-payoff model for stochastic multi-armed bandits, where the payoff (either loss or reward) is delayed for a period that is proportional to the payoff itself. While this captures…
This paper proves the existence of optimal stopping times via elementary functional analytic arguments. The problem is first relaxed into a convex optimization problem over a closed convex subset of the unit ball of the dual of a Banach…
We investigate an infinite-horizon average reward Markov Decision Process (MDP) with delayed, composite, and partially anonymous reward feedback. The delay and compositeness of rewards mean that rewards generated as a result of taking an…
Infinite horizon optimal stopping problems for a L\'evy processes with a two-sided reward function are considered. A two-sided verification theorem is presented in terms of the overall supremum and the overall infimum of the process. A…