English
Related papers

Related papers: Undiscounted optimal stopping with unbounded rewar…

200 papers

We study stochastic linear bandits with heavy-tailed rewards, where the rewards have a finite $(1+\epsilon)$-absolute central moment bounded by $\upsilon$ for some $\epsilon \in (0,1]$. We improve both upper and lower bounds on the minimax…

Machine Learning · Computer Science 2026-01-28 Artin Tajdini , Jonathan Scarlett , Kevin Jamieson

We consider the optimal double stopping time problem defined for each stopping time $S$ by $v(S)=\esssup\{E[\psi(\tau_1, \tau_2) | \F_S], \tau_1, \tau_2 \geq S \}$. Following the optimal one stopping time problem, we study the existence of…

Probability · Mathematics 2009-09-21 Magdalena Kobylanski , Marie-Claire Quenez , Elisabeth Rouy-Mironescu

We study a specific class of finite-horizon mean field optimal stopping problems by means of the dynamic programming approach. In particular, we consider problems where the state process is not affected by the stopping time. Such problems…

Optimization and Control · Mathematics 2025-03-07 Andrea Cosso , Laura Perelli

In the standard models for optimal multiple stopping problems it is assumed that between two exercises there is always a time period of deterministic length $\delta$, the so called refraction period. This prevents the optimal exercise times…

Pricing of Securities · Quantitative Finance 2013-10-17 Sören Christensen , Albrecht Irle , Stephan Jürgens

The Lipschitz bandit problem extends stochastic bandits to a continuous action set defined over a metric space, where the expected reward function satisfies a Lipschitz condition. In this work, we introduce a new problem of Lipschitz bandit…

Machine Learning · Computer Science 2026-02-12 Zhongxuan Liu , Yue Kang , Thomas C. M. Lee

Consider a discrete-time optimal selection problem where one observes a sequence of independent Bernoulli trials and receives a nonnegative reward upon stopping on a success. The aim is to find a single-choice strategy that maximises the…

Probability · Mathematics 2025-12-30 Zakaria Derbazi

We consider the optimal stopping time problem under model uncertainty $R(v)= {\text{ess}\sup\limits}_{ \mathbb{P} \in \mathcal{P}} {\text{ess}\sup\limits}_{\tau \in \mathcal{S}_v} E^\mathbb{P}[Y(\tau) \vert \mathcal{F}_v]$, for every…

Probability · Mathematics 2024-02-23 Ihsan Arharas , Siham Bouhadou , Astrid Hilbert , Youssef Ouknine

We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or abruptly) over time. For this problem setting, we propose an…

Machine Learning · Computer Science 2019-09-11 Pratik Gajane , Ronald Ortner , Peter Auer

We consider the task of estimating a structural model of dynamic decisions by a human agent based upon the observable history of implemented actions and visited states. This problem has an inherent nested structure: in the inner problem, an…

Machine Learning · Computer Science 2024-03-04 Siliang Zeng , Mingyi Hong , Alfredo Garcia

Infinite horizon optimization problems accompany two perplexities. First, the infinite series of utility sequences may diverge. Second, boundary conditions at the infinite terminal time may not be rigorously expressed. In this paper, we…

Optimization and Control · Mathematics 2012-03-20 Dapeng Cai , Gyoshin Nitta

A distributed machine learning platform needs to recruit many heterogeneous worker nodes to finish computation simultaneously. As a result, the overall performance may be degraded due to straggling workers. By introducing redundancy into…

Computer Science and Game Theory · Computer Science 2020-12-17 Ningning Ding , Zhixuan Fang , Lingjie Duan , Jianwei Huang

We consider the optimal stopping problem with non-linear $f$-expectation (induced by a BSDE) without making any regularity assumptions on the reward process $\xi$. and with general filtration. We show that the value family can be aggregated…

Probability · Mathematics 2018-08-02 Miryana Grigorova , Peter Imkeller , Youssef Ouknine , Marie-Claire Quenez

While large reasoning models trained with critic-free reinforcement learning and verifiable rewards (RLVR) represent the state-of-the-art, their practical utility is hampered by ``overthinking'', a critical issue where models generate…

Computation and Language · Computer Science 2026-03-17 Shuyang Jiang , Yusheng Liao , Ya Zhang , Yanfeng Wang , Yu Wang

In this paper, we develop a provably correct optimal control strategy for a finite deterministic transition system. By assuming that penalties with known probabilities of occurrence and dynamics can be sensed locally at the states of the…

Robotics · Computer Science 2013-03-15 Mária Svoreňová , Ivana Černá , Calin Belta

We study a Markov decision problem in which the state space is the set of finite marked point configurations in the plane, the actions represent thinnings, the reward is proportional to the mark sum which is discounted over time, and the…

Probability · Mathematics 2023-09-08 M. N. M. van Lieshout

We consider the problem: is the optimal expected total reward to reach a goal state in a partially observable Markov decision process (POMDP) below a given threshold? We tackle this -- generally undecidable -- problem by computing…

Artificial Intelligence · Computer Science 2022-01-24 Alexander Bork , Joost-Pieter Katoen , Tim Quatmann

Maximum entropy reinforcement learning integrates exploration into policy learning by providing additional intrinsic rewards proportional to the entropy of some distribution. In this paper, we propose a novel approach in which the intrinsic…

Machine Learning · Computer Science 2025-09-30 Adrien Bolland , Gaspard Lambrechts , Damien Ernst

We address an optimal stopping problem over the set of Bermudan-type strategies $\Theta$ (which we understand in a more general sense than the stopping strategies for Bermudan options in finance) and with non-linear operators (non-linear…

Optimization and Control · Mathematics 2023-01-27 Miryana Grigorova , Marie-Claire Quenez , Peng Yuan

In this work we consider optimal stopping problems with conditional convex risk measures called optimised certainty equivalents. Without assuming any kind of time-consistency for the underlying family of risk measures, we derive a novel…

Mathematical Finance · Quantitative Finance 2014-12-16 Denis Belomestny , Volker Kraetschmer

Considering a real-valued diffusion, a real-valued reward function and a positive discount rate, we provide an algorithm to solve the optimal stopping problem consisting in finding the optimal expected discounted reward and the optimal…

Probability · Mathematics 2019-09-24 Fabián Crocce , Ernesto Mordecki