Related papers: Some results on the Gittins index for a normal rew…
A sampling-based method is introduced to approximate the Gittins index for a general family of alternative bandit processes. The approximation consists of a truncation of the optimization horizon and support for the immediate rewards, an…
The Gittins index is a tool that optimally solves a variety of decision-making problems involving uncertainty, including multi-armed bandit problems, minimizing mean latency in queues, and search problems like the Pandora's box model.…
This paper examines a class of singular stochastic control problems with convex objective functions. In Section 2, we use tools from convex analysis to derive necessary and sufficient first order conditions for this class of optimisation…
This paper considers the efficient exact computation of the counterpart of the Gittins index for a finite-horizon discrete-state bandit, which measures for each initial state the average productivity, given by the maximum ratio of expected…
Motivated by global warming issues, we consider a time se- ries that consists of a nondecreasing trend observed with station- ary fluctuations, nonparametric estimation of the trend under monotonicity assumption is considered. The rescaled…
We study dynamic allocation problems for discrete time multi-armed bandits under uncertainty, based on the the theory of nonlinear expectations. We show that, under strong independence of the bandits and with some relaxation in the…
Gini index is a widely used measure of economic inequality. This article develops a general theory for constructing a confidence interval for Gini index with a specified confidence coefficient and a specified width. Fixed sample size…
We prove a general theorem to bound the total variation distance between the distribution of an integer valued random variable of interest and an appropriate discretized normal distribution. We apply the theorem to 2-runs in a sequence of…
Many discrete-time optimal stopping problems are known to have more tractable limit forms based on a planar Poisson process. Using this tool we find a solution to the optimal stopping problem for i.i.d. sequence of $n$ discrete uniform…
We present novel bounds for estimating discrete probability distributions under the $\ell_\infty$ norm. These are nearly optimal in various precise senses, including a kind of instance-optimality. Our data-dependent convergence guarantees…
For a discrete time Markov chain and in line with Strotz' consistent planning we develop a framework for problems of optimal stopping that are time-inconsistent due to the consideration of a non-linear function of an expected reward. We…
This paper presents a new \emph{fast-pivoting} algorithm that computes the $n$ Gittins index values of an $n$-state bandit -- in the discounted and undiscounted cases -- by performing $(2/3) n^3 + O(n^2)$ arithmetic operations, thus…
Designing experiments often requires balancing between learning about the true treatment effects and earning from allocating more samples to the superior treatment. While optimal algorithms for the Multi-Armed Bandit Problem (MABP) provide…
Given a Wiener process with unknown and unobservable drift, we try to estimate this drift as effectively but also as quickly as possible, in the presence of a quadratic penalty for the estimation error and of a fixed, positive cost per unit…
In the budgeted learning problem, we are allowed to experiment on a set of alternatives (given a fixed experimentation budget) with the goal of picking a single alternative with the largest possible expected payoff. Approximation algorithms…
This paper deals with an improvement of the "a-priori stability bounds" on the variation of the action variables and on the stability time obtained from a given Birkhoff normal form around the elliptic equilibrium point of an Hamiltonian…
The dynamic allocation problem, also known as the `multi-armed bandit' problem, simulates a situation in which an agent is faced with a tradeoff between actions that yield an immediate reward and actions whose benefits can only be perceived…
In several recent works on infinite-dimensional systems of ODEs \cite{cao_derivation_2021,cao_explicit_2021,cao_iterative_2024,cao_sticky_2024}, which arise from the mean-field limit of agent-based models in economics and social sciences…
This note gives a short, self-contained, proof of a sharp connection between Gittins indices and Bayesian upper confidence bound algorithms. I consider a Gaussian multi-armed bandit problem with discount factor $\gamma$. The Gittins index…
We obtain upper bounds for the total variation distance between the distributions of two Gibbs point processes in a very general setting. Applications are provided to various well-known processes and settings from spatial statistics and…