English

Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning

Machine Learning 2018-01-03 v3 Artificial Intelligence Machine Learning

Abstract

Statistical performance bounds for reinforcement learning (RL) algorithms can be critical for high-stakes applications like healthcare. This paper introduces a new framework for theoretically measuring the performance of such algorithms called Uniform-PAC, which is a strengthening of the classical Probably Approximately Correct (PAC) framework. In contrast to the PAC framework, the uniform version may be used to derive high probability regret guarantees and so forms a bridge between the two setups that has been missing in the literature. We demonstrate the benefits of the new framework for finite-state episodic MDPs with a new algorithm that is Uniform-PAC and simultaneously achieves optimal regret and PAC guarantees except for a factor of the horizon.

Keywords

Cite

@article{arxiv.1703.07710,
  title  = {Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning},
  author = {Christoph Dann and Tor Lattimore and Emma Brunskill},
  journal= {arXiv preprint arXiv:1703.07710},
  year   = {2018}
}

Comments

appears in Neural Information Processing Systems 2017