English

Revisiting stochastic off-policy action-value gradients

Machine Learning 2017-03-14 v2 Machine Learning

Abstract

Off-policy stochastic actor-critic methods rely on approximating the stochastic policy gradient in order to derive an optimal policy. One may also derive the optimal policy by approximating the action-value gradient. The use of action-value gradients is desirable as policy improvement occurs along the direction of steepest ascent. This has been studied extensively within the context of natural gradient actor-critic algorithms and more recently within the context of deterministic policy gradients. In this paper we briefly discuss the off-policy stochastic counterpart to deterministic action-value gradients, as well as an incremental approach for following the policy gradient in lieu of the natural gradient.

Keywords

Cite

@article{arxiv.1703.02102,
  title  = {Revisiting stochastic off-policy action-value gradients},
  author = {Yemi Okesanjo and Victor Kofia},
  journal= {arXiv preprint arXiv:1703.02102},
  year   = {2017}
}
R2 v1 2026-06-22T18:37:41.580Z