English

Better Optimism By Bayes: Adaptive Planning with Rich Models

Artificial Intelligence 2014-02-11 v1 Machine Learning Machine Learning

Abstract

The computational costs of inference and planning have confined Bayesian model-based reinforcement learning to one of two dismal fates: powerful Bayes-adaptive planning but only for simplistic models, or powerful, Bayesian non-parametric models but using simple, myopic planning strategies such as Thompson sampling. We ask whether it is feasible and truly beneficial to combine rich probabilistic models with a closer approximation to fully Bayesian planning. First, we use a collection of counterexamples to show formal problems with the over-optimism inherent in Thompson sampling. Then we leverage state-of-the-art techniques in efficient Bayes-adaptive planning and non-parametric Bayesian methods to perform qualitatively better than both existing conventional algorithms and Thompson sampling on two contextual bandit-like problems.

Keywords

Cite

@article{arxiv.1402.1958,
  title  = {Better Optimism By Bayes: Adaptive Planning with Rich Models},
  author = {Arthur Guez and David Silver and Peter Dayan},
  journal= {arXiv preprint arXiv:1402.1958},
  year   = {2014}
}

Comments

11 pages, 11 figures

R2 v1 2026-06-22T03:04:20.701Z