Understanding Memory-Regret Trade-Off for Streaming Stochastic Multi-Armed Bandits
Machine Learning
2024-07-09 v2 Data Structures and Algorithms
Machine Learning
Abstract
We study the stochastic multi-armed bandit problem in the -pass streaming model. In this problem, the arms are present in a stream and at most arms and their statistics can be stored in the memory. We give a complete characterization of the optimal regret in terms of and . Specifically, we design an algorithm with regret and complement it with an lower bound when the number of rounds is sufficiently large. Our results are tight up to a logarithmic factor in and .
Keywords
Cite
@article{arxiv.2405.19752,
title = {Understanding Memory-Regret Trade-Off for Streaming Stochastic Multi-Armed Bandits},
author = {Yuchen He and Zichun Ye and Chihao Zhang},
journal= {arXiv preprint arXiv:2405.19752},
year = {2024}
}