English
Related papers

Related papers: Undiscounted optimal stopping with unbounded rewar…

200 papers

We study one-sided and $\alpha$-correct sequential hypothesis testing for data generated by an ergodic Markov chain. The null hypothesis is that the unknown transition matrix belongs to a prescribed set $P$ of stochastic matrices, and the…

Statistics Theory · Mathematics 2026-02-20 Alhad Sethi , Kavali Sofia Sagar , Shubhada Agrawal , Debabrota Basu , P. N. Karthik

We present a method to solve optimal stopping problems in infinite horizon for a L\'evy process when the reward function can be non-monotone. To solve the problem we introduce two new objects. Firstly, we define a random variable $\eta(x)$…

Probability · Mathematics 2015-10-06 Elena Boguslavskaya

The problem of optimal stopping with finite horizon in discrete time is considered in view of maximizing the expected gain. The algorithm proposed in this paper is completely nonparametric in the sense that it uses observed data from the…

Statistics Theory · Mathematics 2013-07-24 Michael Kohler , Harro Walk

We study the properties of the free boundaries and the corresponding hitting times in the context of optimal stopping in discrete time. We first prove the continuity of the map from the boundaries to the expected value of the corresponding…

Probability · Mathematics 2025-04-16 H. Mete Soner , Valentin Tissot-Daguette

In this paper we consider discrete and continuous time risk sensitive optimal stopping problem. Using suitable properties of the underlying Feller-Markov process we prove continuity of the optimal stopping value function and provide formula…

Optimization and Control · Mathematics 2021-03-31 Damian Jelito , Marcin Pitera , Łukasz Stettner

In constrained Markov decision processes (CMDPs) with adversarial rewards and constraints, a well-known impossibility result prevents any algorithm from attaining both sublinear regret and sublinear constraint violation, when competing…

Machine Learning · Computer Science 2024-09-27 Francesco Emanuele Stradi , Anna Lunghi , Matteo Castiglioni , Alberto Marchesi , Nicola Gatti

We study the optimal multiple stopping time problem defined for each stopping time $S$ by $v(S)=\operatorname {ess}\sup_{\tau_1,...,\tau_d\geq S}E[\psi(\tau_1,...,\tau_d)|\mathcal{F}_S]$. The key point is the construction of a new reward…

Probability · Mathematics 2011-08-30 Magdalena Kobylanski , Marie-Claire Quenez , Elisabeth Rouy-Mironescu

The paper addresses the problem of computing maximal conditional expected accumulated rewards until reaching a target state (briefly called maximal conditional expectations) in finite-state Markov decision processes where the condition is…

Logic in Computer Science · Computer Science 2023-03-07 Christel Baier , Joachim Klein , Sascha Klüppelholz , Sascha Wunderlich

For a discrete time Markov chain and in line with Strotz' consistent planning we develop a framework for problems of optimal stopping that are time-inconsistent due to the consideration of a non-linear function of an expected reward. We…

Optimization and Control · Mathematics 2020-01-23 Sören Christensen , Kristoffer Lindensjö

We consider optimal stopping problems, in which a sequence of independent random variables is drawn from a known continuous density. The objective of such problems is to find a procedure which maximizes the expected reward; this is often…

Probability · Mathematics 2020-12-07 Hugh Entwistle , Christopher Lustri , Georgy Sofronov

A general result on the method of randomized stopping is proved. It is applied to optimal stopping of controlled diffusion processes with unbounded coefficients to reduce it to an optimal control problem without stopping. This is motivated…

Probability · Mathematics 2008-05-15 Istvan Gyongy , David Siska

Markov chains are the de facto finite-state model for stochastic dynamical systems, and Markov decision processes (MDPs) extend Markov chains by incorporating non-deterministic behaviors. Given an MDP and rewards on states, a classical…

Logic in Computer Science · Computer Science 2024-11-13 Krishnendu Chatterjee , Laurent Doyen

This paper is devoted to studying constrained continuous-time Markov decision processes (MDPs) in the class of randomized policies depending on state histories. The transition rates may be unbounded, the reward and costs are admitted to be…

Probability · Mathematics 2012-01-04 Xianping Guo , Xinyuan Song

We consider mean-field control problems in discrete time with discounted reward, infinite time horizon and compact state and action space. The existence of optimal policies is shown and the limiting mean-field problem is derived when the…

Optimization and Control · Mathematics 2025-10-16 Nicole Bäuerle

We present a method to find an optimal policy with respect to a reward function for a discounted Markov decision process under general linear temporal logic (LTL) specifications. Previous work has either focused on maximizing a cumulative…

Systems and Control · Electrical Eng. & Systems 2021-03-24 Krishna C. Kalagarla , Rahul Jain , Pierluigi Nuzzo

Standard Markovian optimal stopping problems are consistent in the sense that the first entrance time into the stopping set is optimal for each initial state of the process. Clearly, the usual concept of optimality cannot in a…

Optimization and Control · Mathematics 2018-12-05 Sören Christensen , Kristoffer Lindensjö

Markov reward processes (MRPs) are used to model stochastic phenomena arising in operations research, control engineering, robotics, and artificial intelligence, as well as communication and transportation networks. In many of these cases,…

Machine Learning · Statistics 2020-09-17 Ashwin Pananjady , Martin J. Wainwright

In this paper, we study the optimal stopping problem in the so-called exploratory framework, in which the agent takes actions randomly conditioning on current state and an entropy-regularized term is added to the reward functional. Such a…

Optimization and Control · Mathematics 2023-09-04 Yuchao Dong

Markov decision processes (MDPs) with rewards are a widespread and well-studied model for systems that make both probabilistic and nondeterministic choices. A fundamental result about MDPs is that their minimal and maximal expected rewards…

Logic in Computer Science · Computer Science 2024-11-26 Kevin Batz , Benjamin Lucien Kaminski , Christoph Matheja , Tobias Winkler

We present the first finite time global convergence analysis of policy gradient in the context of infinite horizon average reward Markov decision processes (MDPs). Specifically, we focus on ergodic tabular MDPs with finite state and action…

Machine Learning · Computer Science 2024-03-12 Navdeep Kumar , Yashaswini Murthy , Itai Shufaro , Kfir Y. Levy , R. Srikant , Shie Mannor