English

Thompson Sampling for Gaussian Entropic Risk Bandits

Machine Learning 2021-05-17 v1 Machine Learning

Abstract

The multi-armed bandit (MAB) problem is a ubiquitous decision-making problem that exemplifies exploration-exploitation tradeoff. Standard formulations exclude risk in decision making. Risknotably complicates the basic reward-maximising objectives, in part because there is no universally agreed definition of it. In this paper, we consider an entropic risk (ER) measure and explore the performance of a Thompson sampling-based algorithm ERTS under this risk measure by providing regret bounds for ERTS and corresponding instance dependent lower bounds.

Keywords

Cite

@article{arxiv.2105.06960,
  title  = {Thompson Sampling for Gaussian Entropic Risk Bandits},
  author = {Ming Liang Ang and Eloise Y. Y. Lim and Joel Q. L. Chang},
  journal= {arXiv preprint arXiv:2105.06960},
  year   = {2021}
}

Comments

arXiv admin note: text overlap with arXiv:2011.08046

R2 v1 2026-06-24T02:07:27.053Z