English

Convergence of Q-value in case of Gaussian rewards

Optimization and Control 2021-09-14 v1 Machine Learning Machine Learning

Abstract

In this paper, as a study of reinforcement learning, we converge the Q function to unbounded rewards such as Gaussian distribution. From the central limit theorem, in some real-world applications it is natural to assume that rewards follow a Gaussian distribution , but existing proofs cannot guarantee convergence of the Q-function. Furthermore, in the distribution-type reinforcement learning and Bayesian reinforcement learning that have become popular in recent years, it is better to allow the reward to have a Gaussian distribution. Therefore, in this paper, we prove the convergence of the Q-function under the condition of E[r(s,a)2]<E[r(s,a)^2]<\infty, which is much more relaxed than the existing research. Finally, as a bonus, a proof of the policy gradient theorem for distributed reinforcement learning is also posted.

Keywords

Cite

@article{arxiv.2003.03526,
  title  = {Convergence of Q-value in case of Gaussian rewards},
  author = {Konatsu Miyamoto and Masaya Suzuki and Yuma Kigami and Kodai Satake},
  journal= {arXiv preprint arXiv:2003.03526},
  year   = {2021}
}

Comments

10 pages