English

On Minimax Optimal Offline Policy Evaluation

Artificial Intelligence 2014-09-15 v1

Abstract

This paper studies the off-policy evaluation problem, where one aims to estimate the value of a target policy based on a sample of observations collected by another policy. We first consider the multi-armed bandit case, establish a minimax risk lower bound, and analyze the risk of two standard estimators. It is shown, and verified in simulation, that one is minimax optimal up to a constant, while another can be arbitrarily worse, despite its empirical success and popularity. The results are applied to related problems in contextual bandits and fixed-horizon Markov decision processes, and are also related to semi-supervised learning.

Keywords

Cite

@article{arxiv.1409.3653,
  title  = {On Minimax Optimal Offline Policy Evaluation},
  author = {Lihong Li and Remi Munos and Csaba Szepesvari},
  journal= {arXiv preprint arXiv:1409.3653},
  year   = {2014}
}
R2 v1 2026-06-22T05:55:06.075Z