Classical Policy Gradient: Preserving Bellman's Principle of Optimality
Machine Learning
2019-06-10 v1 Machine Learning
Abstract
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.
Cite
@article{arxiv.1906.03063,
title = {Classical Policy Gradient: Preserving Bellman's Principle of Optimality},
author = {Philip S. Thomas and Scott M. Jordan and Yash Chandak and Chris Nota and James Kostas},
journal= {arXiv preprint arXiv:1906.03063},
year = {2019}
}
Comments
1 page, 0 figures