English

Causal Bandits: Learning Good Interventions via Causal Inference

Machine Learning 2016-06-13 v1 Machine Learning

Abstract

We study the problem of using causal models to improve the rate at which good interventions can be learned online in a stochastic environment. Our formalism combines multi-arm bandits and causal inference to model a novel type of bandit feedback that is not exploited by existing approaches. We propose a new algorithm that exploits the causal feedback and prove a bound on its simple regret that is strictly better (in all quantities) than algorithms that do not use the additional causal information.

Keywords

Cite

@article{arxiv.1606.03203,
  title  = {Causal Bandits: Learning Good Interventions via Causal Inference},
  author = {Finnian Lattimore and Tor Lattimore and Mark D. Reid},
  journal= {arXiv preprint arXiv:1606.03203},
  year   = {2016}
}
R2 v1 2026-06-22T14:22:17.965Z