Related papers: A Proof of the Bomber Problem's Spend-It-All Conje…
We study a novel pure exploration problem: the $\epsilon$-Thresholding Bandit Problem (TBP) with fixed confidence in stochastic linear bandits. We prove a lower bound for the sample complexity and extend an algorithm designed for Best Arm…
In this paper, we present approximation algorithms for combinatorial optimization problems under probabilistic constraints. Specifically, we focus on stochastic variants of two important combinatorial optimization problems: the k-center…
The objective of this work is to study continuous-time Markov decision processes on a general Borel state space with both impulsive and continuous controls for the infinite-time horizon discounted cost. The continuous-time controlled…
We consider optimal stopping problems for a Brownian motion and a geometric Brownian motion with a "disorder", assuming that the moment of a disorder is uniformly distributed on a finite interval. Optimal stopping rules are found as the…
We study a version of the classical Cayley-Moser optimal stopping problem, in which a seller must sell an asset by a given deadline, with the offers, which are independent random variables with a known distribution, arriving at random…
We present a methodology for obtaining explicit solutions to infinite time horizon optimal stopping problems involving general, one-dimensional, It\^o diffusions, payoff functions that need not be smooth and state-dependent discounting.…
We consider a continuous time two-armed bandit problem in which incomes are described by Poissonian processes. We develop Bayesian approach with arbitrary prior distribution. We present two versions of recursive equation for determination…
The pinwheel problem is a real-time scheduling problem that asks, given $n$ tasks with periods $a_i \in \mathbb{N}$, whether it is possible to infinitely schedule the tasks, one per time unit, such that every task $i$ is scheduled in every…
In the multiarmed bandit problem a gambler chooses an arm of a slot machine to pull considering a tradeoff between exploration and exploitation. We study the stochastic bandit problem where each arm has a reward distribution supported in a…
Solutions in the form of series expansion, as the Born approximation, are very useful for describing time-independent scattering of quantum particles. In this work, it is mathematically demonstred that such solutions, when applied to…
We consider a singular stochastic control problem, which is called the Monotone Follower Stochastic Control Problem and give sufficient conditions for the existence and uniqueness of a local-time type optimal control. To establish this…
We address the problem of identifying the optimal policy with a fixed confidence level in a multi-armed bandit setup, when \emph{the arms are subject to linear constraints}. Unlike the standard best-arm identification problem which is well…
We consider the problem of adaptively placing sensors along an interval to detect stochastically-generated events. We present a new formulation of the problem as a continuum-armed bandit problem with feedback in the form of partial…
In this paper we consider the contextual multi-armed bandit problem for linear payoffs under a risk-averse criterion. At each round, contexts are revealed for each arm, and the decision maker chooses one arm to pull and receives the…
Given a vector of probability distributions, or arms, each of which can be sampled independently, we consider the problem of identifying the partition to which this vector belongs from a finitely partitioned universe of such vector of…
This note considers a variation of the full-information secretary problem where the random variables to be observed are independent and identically distributed. Consider $X_1,\dots,X_n$ to be an independent sequence of random variables, let…
In this paper we study a representation problem first considered in a simpler version by Bank and El Karoui [2004]. A key ingredient to this problem is a random measure $\mu$ on the time axis which in the present paper is allowed to have…
The pure-exploration problem in stochastic multi-armed bandits aims to find one or more arms with the largest (or near largest) means. Examples include finding an {\epsilon}-good arm, best-arm identification, top-k arm identification, and…
In this paper, we investigate the robust optimal reinsurance,investment,and internal surplus distribution (i.e., consumption) problem for an insurer with Epstein-Zin recursive preferences in an incomplete market. It is assumed that the…
We consider bandit problems involving a large (possibly infinite) collection of arms, in which the expected reward of each arm is a linear function of an $r$-dimensional random vector $\mathbf{Z} \in \mathbb{R}^r$, where $r \geq 2$. The…