Related papers: Undiscounted optimal stopping with unbounded rewar…
We study stochastic linear bandits with heavy-tailed rewards, where the rewards have a finite $(1+\epsilon)$-absolute central moment bounded by $\upsilon$ for some $\epsilon \in (0,1]$. We improve both upper and lower bounds on the minimax…
We consider the optimal double stopping time problem defined for each stopping time $S$ by $v(S)=\esssup\{E[\psi(\tau_1, \tau_2) | \F_S], \tau_1, \tau_2 \geq S \}$. Following the optimal one stopping time problem, we study the existence of…
We study a specific class of finite-horizon mean field optimal stopping problems by means of the dynamic programming approach. In particular, we consider problems where the state process is not affected by the stopping time. Such problems…
In the standard models for optimal multiple stopping problems it is assumed that between two exercises there is always a time period of deterministic length $\delta$, the so called refraction period. This prevents the optimal exercise times…
The Lipschitz bandit problem extends stochastic bandits to a continuous action set defined over a metric space, where the expected reward function satisfies a Lipschitz condition. In this work, we introduce a new problem of Lipschitz bandit…
Consider a discrete-time optimal selection problem where one observes a sequence of independent Bernoulli trials and receives a nonnegative reward upon stopping on a success. The aim is to find a single-choice strategy that maximises the…
We consider the optimal stopping time problem under model uncertainty $R(v)= {\text{ess}\sup\limits}_{ \mathbb{P} \in \mathcal{P}} {\text{ess}\sup\limits}_{\tau \in \mathcal{S}_v} E^\mathbb{P}[Y(\tau) \vert \mathcal{F}_v]$, for every…
We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or abruptly) over time. For this problem setting, we propose an…
We consider the task of estimating a structural model of dynamic decisions by a human agent based upon the observable history of implemented actions and visited states. This problem has an inherent nested structure: in the inner problem, an…
Infinite horizon optimization problems accompany two perplexities. First, the infinite series of utility sequences may diverge. Second, boundary conditions at the infinite terminal time may not be rigorously expressed. In this paper, we…
A distributed machine learning platform needs to recruit many heterogeneous worker nodes to finish computation simultaneously. As a result, the overall performance may be degraded due to straggling workers. By introducing redundancy into…
We consider the optimal stopping problem with non-linear $f$-expectation (induced by a BSDE) without making any regularity assumptions on the reward process $\xi$. and with general filtration. We show that the value family can be aggregated…
While large reasoning models trained with critic-free reinforcement learning and verifiable rewards (RLVR) represent the state-of-the-art, their practical utility is hampered by ``overthinking'', a critical issue where models generate…
In this paper, we develop a provably correct optimal control strategy for a finite deterministic transition system. By assuming that penalties with known probabilities of occurrence and dynamics can be sensed locally at the states of the…
We study a Markov decision problem in which the state space is the set of finite marked point configurations in the plane, the actions represent thinnings, the reward is proportional to the mark sum which is discounted over time, and the…
We consider the problem: is the optimal expected total reward to reach a goal state in a partially observable Markov decision process (POMDP) below a given threshold? We tackle this -- generally undecidable -- problem by computing…
Maximum entropy reinforcement learning integrates exploration into policy learning by providing additional intrinsic rewards proportional to the entropy of some distribution. In this paper, we propose a novel approach in which the intrinsic…
We address an optimal stopping problem over the set of Bermudan-type strategies $\Theta$ (which we understand in a more general sense than the stopping strategies for Bermudan options in finance) and with non-linear operators (non-linear…
In this work we consider optimal stopping problems with conditional convex risk measures called optimised certainty equivalents. Without assuming any kind of time-consistency for the underlying family of risk measures, we derive a novel…
Considering a real-valued diffusion, a real-valued reward function and a positive discount rate, we provide an algorithm to solve the optimal stopping problem consisting in finding the optimal expected discounted reward and the optimal…