English

PAC-Bayesian Lifelong Learning For Multi-Armed Bandits

Machine Learning 2022-03-17 v1 Machine Learning

Abstract

We present a PAC-Bayesian analysis of lifelong learning. In the lifelong learning problem, a sequence of learning tasks is observed one-at-a-time, and the goal is to transfer information acquired from previous tasks to new learning tasks. We consider the case when each learning task is a multi-armed bandit problem. We derive lower bounds on the expected average reward that would be obtained if a given multi-armed bandit algorithm was run in a new task with a particular prior and for a set number of steps. We propose lifelong learning algorithms that use our new bounds as learning objectives. Our proposed algorithms are evaluated in several lifelong multi-armed bandit problems and are found to perform better than a baseline method that does not use generalisation bounds.

Keywords

Cite

@article{arxiv.2203.03303,
  title  = {PAC-Bayesian Lifelong Learning For Multi-Armed Bandits},
  author = {Hamish Flynn and David Reeb and Melih Kandemir and Jan Peters},
  journal= {arXiv preprint arXiv:2203.03303},
  year   = {2022}
}

Comments

29 pages, 5 figures

R2 v1 2026-06-24T10:04:23.115Z