The ODE Method for Asymptotic Statistics in Stochastic Approximation and Reinforcement Learning
Abstract
The paper concerns the -dimensional stochastic approximation recursion, where is a stochastic process on a general state space, satisfying a conditional Markov property that allows for parameter-dependent noise. The main results are established under additional conditions on the mean flow and a version of the Donsker-Varadhan Lyapunov drift condition known as (DV3): (i) An appropriate Lyapunov function is constructed that implies convergence of the estimates in . (ii) A functional central limit theorem (CLT) is established, as well as the usual one-dimensional CLT for the normalized error. Moment bounds combined with the CLT imply convergence of the normalized covariance to the asymptotic covariance in the CLT, where . (iii) The CLT holds for the normalized version , of the averaged parameters , subject to standard assumptions on the step-size. Moreover, the covariance in the CLT coincides with the minimal covariance of Polyak and Ruppert. (iv) An example is given where and are linear in , and is a geometrically ergodic Markov chain but does not satisfy (DV3). While the algorithm is convergent, the second moment of is unbounded and in fact diverges. This arXiv version represents a major extension of the results in prior versions.The main results now allow for parameter-dependent noise, as is often the case in applications to reinforcement learning.
Cite
@article{arxiv.2110.14427,
title = {The ODE Method for Asymptotic Statistics in Stochastic Approximation and Reinforcement Learning},
author = {Vivek Borkar and Shuhang Chen and Adithya Devraj and Ioannis Kontoyiannis and Sean Meyn},
journal= {arXiv preprint arXiv:2110.14427},
year = {2024}
}
Comments
This arXiv version represents a major extension of the results in prior versions.The main results now allow for parameter-dependent noise, as is often the case in applications to reinforcement learning. 2 figures