English

Near-Optimal MNL Bandits Under Risk Criteria

Machine Learning 2021-03-17 v3 Machine Learning

Abstract

We study MNL bandits, which is a variant of the traditional multi-armed bandit problem, under risk criteria. Unlike the ordinary expected revenue, risk criteria are more general goals widely used in industries and bussiness. We design algorithms for a broad class of risk criteria, including but not limited to the well-known conditional value-at-risk, Sharpe ratio and entropy risk, and prove that they suffer a near-optimal regret. As a complement, we also conduct experiments with both synthetic and real data to show the empirical performance of our proposed algorithms.

Keywords

Cite

@article{arxiv.2009.12511,
  title  = {Near-Optimal MNL Bandits Under Risk Criteria},
  author = {Guangyu Xi and Chao Tao and Yuan Zhou},
  journal= {arXiv preprint arXiv:2009.12511},
  year   = {2021}
}

Comments

AAAI2021