English

A Fast Convergence Theory for Offline Decision Making

Machine Learning 2024-12-04 v2 Machine Learning

Abstract

This paper proposes the first generic fast convergence result in general function approximation for offline decision making problems, which include offline reinforcement learning (RL) and off-policy evaluation (OPE) as special cases. To unify different settings, we introduce a framework called Decision Making with Offline Feedback (DMOF), which captures a wide range of offline decision making problems. Within this framework, we propose a simple yet powerful algorithm called Empirical Decision with Divergence (EDD), whose upper bound can be termed as a coefficient named Empirical Offline Estimation Coefficient (EOEC). We show that EOEC is instance-dependent and actually measures the correlation of the problem. When assuming partial coverage in the dataset, EOEC will reduce in a rate of 1/N1/N where NN is the size of the dataset, endowing EDD with a fast convergence guarantee. Finally, we complement the above results with a lower bound in the DMOF framework, which further demonstrates the soundness of our theory.

Keywords

Cite

@article{arxiv.2406.01378,
  title  = {A Fast Convergence Theory for Offline Decision Making},
  author = {Chenjie Mao and Qiaosheng Zhang},
  journal= {arXiv preprint arXiv:2406.01378},
  year   = {2024}
}
R2 v1 2026-06-28T16:51:13.740Z