Related papers: Delayed Gambler's Ruin
Consider planning a trip in a train network. In contrast to, say, a road network, the edges are temporal, i.e., they are only available at certain times. Another important difficulty is that trains, unfortunately, sometimes get delayed.…
We address in this paper the approximation problem of distributed delays. Such elements are convolution operators with kernel having bounded support, and appear in the control of time-delay systems. From the rich literature on this topic,…
The fidelity bandits problem is a variant of the $K$-armed bandit problem in which the reward of each arm is augmented by a fidelity reward that provides the player with an additional payoff depending on how 'loyal' the player has been to…
There are many algorithms for regret minimisation in episodic reinforcement learning. This problem is well-understood from a theoretical perspective, providing that the sequences of states, actions and rewards associated with each episode…
I model a rational agent who experiences endogenous deadline pressure in the face of a fixed future deadline. The agent holds a resource stock, and opportunities to spend resources arise randomly according to a Poisson process. When the…
We introduce the hybrid risk process, constructed via a time-change transformation applied to the solution of a hybrid stochastic differential equation. The framework covers several modern ruin settings, incorporating features like…
We study the problem of minimising regret in two-armed bandit problems with Gaussian rewards. Our objective is to use this simple setting to illustrate that strategies based on an exploration phase (up to a stopping time) followed by…
The aim of this paper is to construct the confidence interval of the ultimate ruin probability under the insurance surplus driven by a L\'evy process. Assuming a parametric family for the L\'evy measures, we estimate the parameter from the…
This paper deals with the discrete-time risk model with nonidentically distributed claims. We suppose that the claims repeat with time periods of three units, that is, claim distributions coincide at times $\{1,4,7,\ldots\}$, at times…
Following an article by Muller and Pflug, we study the adjustment coefficient of ruin theory in a context of temporal dependency. We provide a consistent estimator of this coefficient, and perform some simulations.
No matter how much some gamblers occasionally win, as long as they continue to gamble, sooner or later they will lose more to the casino, which is the so-called long bet will lose. Our results demonstrate the counter-intuitive phenomenon,…
This paper studies a new and more general axiomatization than one presented previously for preference on likelihood gambles. Likelihood gambles describe actions in a situation where a decision maker knows multiple probabilistic models and a…
Many poker systems, whether created with heuristics or machine learning, rely on the probability of winning as a key input. However calculating the precise probability using combinatorics is an intractable problem, so instead we approximate…
We study a new type of K-armed bandit problem where the expected return of one arm may depend on the returns of other arms. We present a new algorithm for this general class of problems and show that under certain circumstances it is…
We consider decentralized restless multi-armed bandit problems with unknown dynamics and multiple players. The reward state of each arm transits according to an unknown Markovian rule when it is played and evolves according to an arbitrary…
Motivated by applications to online advertising and recommender systems, we consider a game-theoretic model with delayed rewards and asynchronous, payoff-based feedback. In contrast to previous work on delayed multi-armed bandits, we focus…
This paper develops asymptotics and approximations for ruin probabilities in a multivariate risk setting. We consider a model in which the individual reserve processes are driven by a common Markovian environmental process. We subsequently…
We consider how an agent should update her uncertainty when it is represented by a set $\P$ of probability distributions and the agent observes that a random variable $X$ takes on value $x$, given that the agent makes decisions using the…
We study the stochastic Multi-Armed Bandit (MAB) problem with random delays in the feedback received by the algorithm. We consider two settings: the reward-dependent delay setting, where realized delays may depend on the stochastic rewards,…
We study the gambler's ruin problem for the Elephant Random Walk, focusing on escape time from a symmetric interval of the form $\{-N, \ldots, N\}$. As our main result, we derive tight exponential bounds for the tail of this escape time. We…