English

Variance-Reduced Off-Policy TDC Learning: Non-Asymptotic Convergence Analysis

Machine Learning 2023-05-23 v4 Optimization and Control

Abstract

Variance reduction techniques have been successfully applied to temporal-difference (TD) learning and help to improve the sample complexity in policy evaluation. However, the existing work applied variance reduction to either the less popular one time-scale TD algorithm or the two time-scale GTD algorithm but with a finite number of i.i.d.\ samples, and both algorithms apply to only the on-policy setting. In this work, we develop a variance reduction scheme for the two time-scale TDC algorithm in the off-policy setting and analyze its non-asymptotic convergence rate over both i.i.d.\ and Markovian samples. In the i.i.d.\ setting, our algorithm {matches the best-known lower bound O~(ϵ1\tilde{O}(\epsilon^{-1}).} In the Markovian setting, our algorithm achieves the state-of-the-art sample complexity O(ϵ1logϵ1)O(\epsilon^{-1} \log {\epsilon}^{-1}) that is near-optimal. Experiments demonstrate that the proposed variance-reduced TDC achieves a smaller asymptotic convergence error than both the conventional TDC and the variance-reduced TD.

Keywords

Cite

@article{arxiv.2010.13272,
  title  = {Variance-Reduced Off-Policy TDC Learning: Non-Asymptotic Convergence Analysis},
  author = {Shaocong Ma and Yi Zhou and Shaofeng Zou},
  journal= {arXiv preprint arXiv:2010.13272},
  year   = {2023}
}

Comments

Accepted for publication in NeurIPS 2020

R2 v1 2026-06-23T19:38:19.525Z