English

Online Reinforcement Learning for Periodic MDP

Machine Learning 2022-07-26 v1

Abstract

We study learning in periodic Markov Decision Process(MDP), a special type of non-stationary MDP where both the state transition probabilities and reward functions vary periodically, under the average reward maximization setting. We formulate the problem as a stationary MDP by augmenting the state space with the period index, and propose a periodic upper confidence bound reinforcement learning-2 (PUCRL2) algorithm. We show that the regret of PUCRL2 varies linearly with the period and as sub-linear with the horizon length. Numerical results demonstrate the efficacy of PUCRL2.

Keywords

Cite

@article{arxiv.2207.12045,
  title  = {Online Reinforcement Learning for Periodic MDP},
  author = {Ayush Aniket and Arpan Chattopadhyay},
  journal= {arXiv preprint arXiv:2207.12045},
  year   = {2022}
}
R2 v1 2026-06-25T01:11:50.825Z