A Note on a Tight Lower Bound for MNL-Bandit Assortment Selection Models
Machine Learning
2018-10-01 v3 Machine Learning
Abstract
In this short note we consider a dynamic assortment planning problem under the capacitated multinomial logit (MNL) bandit model. We prove a tight lower bound on the accumulated regret that matches existing regret upper bounds for all parameters (time horizon , number of items and maximum assortment capacity ) up to logarithmic factors. Our results close an gap between upper and lower regret bounds from existing works.
Keywords
Cite
@article{arxiv.1709.06109,
title = {A Note on a Tight Lower Bound for MNL-Bandit Assortment Selection Models},
author = {Xi Chen and Yining Wang},
journal= {arXiv preprint arXiv:1709.06109},
year = {2018}
}
Comments
Final version, 4 pages (double column)