Skyline Identification in Multi-Armed Bandits
Abstract
We introduce a variant of the classical PAC multi-armed bandit problem. There is an ordered set of arms , each with some stochastic reward drawn from some unknown bounded distribution. The goal is to identify the of the set , consisting of all arms such that has larger expected reward than all lower-numbered arms . We define a natural notion of an -approximate skyline and prove matching upper and lower bounds for identifying an -skyline. Specifically, we show that in order to identify an -skyline from among arms with probability , samples are necessary and sufficient. When , our results improve over the naive algorithm, which draws enough samples to approximate the expected reward of every arm; the algorithm of (Auer et al., AISTATS'16) for Pareto-optimal arm identification is likewise superseded. Our results show that the sample complexity of the skyline problem lies strictly in between that of best arm identification (Even-Dar et al., COLT'02) and that of approximating the expected reward of every arm.
Keywords
Cite
@article{arxiv.1711.04213,
title = {Skyline Identification in Multi-Armed Bandits},
author = {Albert Cheu and Ravi Sundaram and Jonathan Ullman},
journal= {arXiv preprint arXiv:1711.04213},
year = {2018}
}
Comments
18 pages, 2 Figures; an ALT'18/ISIT'18 submission