English

RL for Latent MDPs: Regret Guarantees and a Lower Bound

Machine Learning 2021-02-10 v1

Abstract

In this work, we consider the regret minimization problem for reinforcement learning in latent Markov Decision Processes (LMDP). In an LMDP, an MDP is randomly drawn from a set of MM possible MDPs at the beginning of the interaction, but the identity of the chosen MDP is not revealed to the agent. We first show that a general instance of LMDPs requires at least Ω((SA)M)\Omega((SA)^M) episodes to even approximate the optimal policy. Then, we consider sufficient assumptions under which learning good policies requires polynomial number of episodes. We show that the key link is a notion of separation between the MDP system dynamics. With sufficient separation, we provide an efficient algorithm with local guarantee, {\it i.e.,} providing a sublinear regret guarantee when we are given a good initialization. Finally, if we are given standard statistical sufficiency assumptions common in the Predictive State Representation (PSR) literature (e.g., Boots et al.) and a reachability assumption, we show that the need for initialization can be removed.

Keywords

Cite

@article{arxiv.2102.04939,
  title  = {RL for Latent MDPs: Regret Guarantees and a Lower Bound},
  author = {Jeongyeol Kwon and Yonathan Efroni and Constantine Caramanis and Shie Mannor},
  journal= {arXiv preprint arXiv:2102.04939},
  year   = {2021}
}
R2 v1 2026-06-23T22:59:15.598Z