English

Thompson Sampling on Asymmetric $\alpha$-Stable Bandits

Machine Learning 2022-03-28 v2 Machine Learning

Abstract

In algorithm optimization in reinforcement learning, how to deal with the exploration-exploitation dilemma is particularly important. Multi-armed bandit problem can optimize the proposed solutions by changing the reward distribution to realize the dynamic balance between exploration and exploitation. Thompson Sampling is a common method for solving multi-armed bandit problem and has been used to explore data that conform to various laws. In this paper, we consider the Thompson Sampling approach for multi-armed bandit problem, in which rewards conform to unknown asymmetric α\alpha-stable distributions and explore their applications in modelling financial and wireless data.

Keywords

Cite

@article{arxiv.2203.10214,
  title  = {Thompson Sampling on Asymmetric $\alpha$-Stable Bandits},
  author = {Zhendong Shi and Ercan E. Kuruoglu and Xiaoli Wei},
  journal= {arXiv preprint arXiv:2203.10214},
  year   = {2022}
}

Comments

8 pages, 4 figures

R2 v1 2026-06-24T10:18:56.199Z