English

Classical Policy Gradient: Preserving Bellman's Principle of Optimality

Machine Learning 2019-06-10 v1 Machine Learning

Abstract

We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.

Cite

@article{arxiv.1906.03063,
  title  = {Classical Policy Gradient: Preserving Bellman's Principle of Optimality},
  author = {Philip S. Thomas and Scott M. Jordan and Yash Chandak and Chris Nota and James Kostas},
  journal= {arXiv preprint arXiv:1906.03063},
  year   = {2019}
}

Comments

1 page, 0 figures

R2 v1 2026-06-23T09:46:57.555Z