Concentration of Contractive Stochastic Approximation and Reinforcement Learning
Machine Learning
2022-06-14 v4 Systems and Control
Systems and Control
Abstract
Using a martingale concentration inequality, concentration bounds `from time on' are derived for stochastic approximation algorithms with contractive maps and both martingale difference and Markov noises. These are applied to reinforcement learning algorithms, in particular to asynchronous Q-learning and TD(0).
Cite
@article{arxiv.2106.14308,
title = {Concentration of Contractive Stochastic Approximation and Reinforcement Learning},
author = {Siddharth Chandak and Vivek S. Borkar and Parth Dodhia},
journal= {arXiv preprint arXiv:2106.14308},
year = {2022}
}
Comments
20 pages, Accepted for publication in Stochastic Systems