English

The ODE Method for Asymptotic Statistics in Stochastic Approximation and Reinforcement Learning

Statistics Theory 2024-11-18 v6 Machine Learning Statistics Theory

Abstract

The paper concerns the dd-dimensional stochastic approximation recursion, θn+1=θn+αn+1f(θn,Φn+1) \theta_{n+1}= \theta_n + \alpha_{n + 1} f(\theta_n, \Phi_{n+1}) where {Φn} \{ \Phi_n \} is a stochastic process on a general state space, satisfying a conditional Markov property that allows for parameter-dependent noise. The main results are established under additional conditions on the mean flow and a version of the Donsker-Varadhan Lyapunov drift condition known as (DV3): (i) An appropriate Lyapunov function is constructed that implies convergence of the estimates in L4L_4. (ii) A functional central limit theorem (CLT) is established, as well as the usual one-dimensional CLT for the normalized error. Moment bounds combined with the CLT imply convergence of the normalized covariance E[znznT]\textsf{E}[ z_n z_n^T ] to the asymptotic covariance in the CLT, where zn=:(θnθ)/αnz_n =: (\theta_n-\theta^*)/\sqrt{\alpha_n}. (iii) The CLT holds for the normalized version znPR=:n[θnPRθ]z^{\text{PR}}_n =: \sqrt{n} [\theta^{\text{PR}}_n -\theta^*], of the averaged parameters θnPR=:n1k=1nθk\theta^{\text{PR}}_n =:n^{-1} \sum_{k=1}^n\theta_k, subject to standard assumptions on the step-size. Moreover, the covariance in the CLT coincides with the minimal covariance of Polyak and Ruppert. (iv) An example is given where ff and fˉ\bar{f} are linear in θ\theta, and Φ\Phi is a geometrically ergodic Markov chain but does not satisfy (DV3). While the algorithm is convergent, the second moment of θn\theta_n is unbounded and in fact diverges. This arXiv version represents a major extension of the results in prior versions.The main results now allow for parameter-dependent noise, as is often the case in applications to reinforcement learning.

Keywords

Cite

@article{arxiv.2110.14427,
  title  = {The ODE Method for Asymptotic Statistics in Stochastic Approximation and Reinforcement Learning},
  author = {Vivek Borkar and Shuhang Chen and Adithya Devraj and Ioannis Kontoyiannis and Sean Meyn},
  journal= {arXiv preprint arXiv:2110.14427},
  year   = {2024}
}

Comments

This arXiv version represents a major extension of the results in prior versions.The main results now allow for parameter-dependent noise, as is often the case in applications to reinforcement learning. 2 figures

R2 v1 2026-06-24T07:14:01.750Z