English

A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting

Machine Learning 2020-11-05 v2 Artificial Intelligence Machine Learning

Abstract

Recently, Wang et al. (2020) showed a highly intriguing hardness result for batch reinforcement learning (RL) with linearly realizable value function and good feature coverage in the finite-horizon case. In this note we show that once adapted to the discounted setting, the construction can be simplified to a 2-state MDP with 1-dimensional features, such that learning is impossible even with an infinite amount of data.

Keywords

Cite

@article{arxiv.2011.01075,
  title  = {A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting},
  author = {Philip Amortila and Nan Jiang and Tengyang Xie},
  journal= {arXiv preprint arXiv:2011.01075},
  year   = {2020}
}