English

Finite-Time Accuracy of Temporal-Difference Learning Under Schur-Stable Recursions

Machine Learning 2026-02-02 v7 Systems and Control Systems and Control

Abstract

Temporal difference (TD) learning is a cornerstone reinforcement learning (RL) method for policy evaluation, where the goal is to estimate the value function of a Markov decision process under a fixed policy. While a substantial body of work has established its convergence and stability properties, more recent efforts have focused on its statistical efficiency through finite-time error bounds. In this paper, we advance this line of research by developing a new finite-time error analysis for tabular TD learning that directly exploits a discrete-time stochastic linear system representation and leverages Schur stability of the associated matrices. Beyond the specific bounds obtained, the proposed framework provides a reusable template for analyzing TD learning and related RL algorithms, and it offers control-theoretic insights that may guide future developments in finite-sample RL theory.

Keywords

Cite

@article{arxiv.2204.10479,
  title  = {Finite-Time Accuracy of Temporal-Difference Learning Under Schur-Stable Recursions},
  author = {Donghwan Lee and Do Wan Kim},
  journal= {arXiv preprint arXiv:2204.10479},
  year   = {2026}
}

Comments

arXiv admin note: text overlap with arXiv:2112.14417

R2 v1 2026-06-24T10:55:28.787Z