English

On Oracle-Efficient PAC RL with Rich Observations

Machine Learning 2019-01-18 v4 Machine Learning

Abstract

We study the computational tractability of PAC reinforcement learning with rich observations. We present new provably sample-efficient algorithms for environments with deterministic hidden state dynamics and stochastic rich observations. These methods operate in an oracle model of computation -- accessing policy and value function classes exclusively through standard optimization primitives -- and therefore represent computationally efficient alternatives to prior algorithms that require enumeration. With stochastic hidden state dynamics, we prove that the only known sample-efficient algorithm, OLIVE, cannot be implemented in the oracle model. We also present several examples that illustrate fundamental challenges of tractable PAC reinforcement learning in such general settings.

Keywords

Cite

@article{arxiv.1803.00606,
  title  = {On Oracle-Efficient PAC RL with Rich Observations},
  author = {Christoph Dann and Nan Jiang and Akshay Krishnamurthy and Alekh Agarwal and John Langford and Robert E. Schapire},
  journal= {arXiv preprint arXiv:1803.00606},
  year   = {2019}
}

Comments

appeared at NeurIPS 18; full paper including appendix; updated style file

R2 v1 2026-06-23T00:38:44.222Z