Differential TD Learning for Value Function Approximation

Adithya M. Devraj; Sean P. Meyn

Differential TD Learning for Value Function Approximation

Systems and Control 2018-12-27 v3 Machine Learning Optimization and Control

Authors: Adithya M. Devraj , Sean P. Meyn

Abstract

Value functions arise as a component of algorithms as well as performance metrics in statistics and engineering applications. Computation of the associated Bellman equations is numerically challenging in all but a few special cases. A popular approximation technique is known as Temporal Difference (TD) learning. The algorithm introduced in this paper is intended to resolve two well-known problems with this approach: In the discounted-cost setting, the variance of the algorithm diverges as the discount factor approaches unity. Second, for the average cost setting, unbiased algorithms exist only in special cases. It is shown that the gradient of any of these value functions admits a representation that lends itself to algorithm design. Based on this result, the new differential TD method is obtained for Markovian models on Euclidean space with smooth dynamics. Numerical examples show remarkable improvements in performance. In application to speed scaling, variance is reduced by two orders of magnitude.

Keywords

dimensionality reduction optimization stochastic gradient descent

Cite

@article{arxiv.1604.01828,
  title  = {Differential TD Learning for Value Function Approximation},
  author = {Adithya M. Devraj and Sean P. Meyn},
  journal= {arXiv preprint arXiv:1604.01828},
  year   = {2018}
}

Differential TD Learning for Value Function Approximation

Abstract

Keywords

Cite

Related papers