English

Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs

Machine Learning 2024-03-12 v3

Abstract

Reinforcement learning generalizes multi-armed bandit problems with additional difficulties of a longer planning horizon and unknown transition kernel. We explore a black-box reduction from discounted infinite-horizon tabular reinforcement learning to multi-armed bandits, where, specifically, an independent bandit learner is placed in each state. We show that, under ergodicity and fast mixing assumptions, any slowly changing adversarial bandit algorithm achieving optimal regret in the adversarial bandit setting can also attain optimal expected regret in infinite-horizon discounted Markov decision processes, with respect to the number of rounds TT. Furthermore, we examine our reduction using a specific instance of the exponential-weight algorithm.

Keywords

Cite

@article{arxiv.2205.09056,
  title  = {Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs},
  author = {Ian A. Kash and Lev Reyzin and Zishun Yu},
  journal= {arXiv preprint arXiv:2205.09056},
  year   = {2024}
}

Comments

ALT 24

R2 v1 2026-06-24T11:21:20.256Z