Near-Optimal Regret for Adversarial MDP with Delayed Bandit Feedback
Machine Learning
2023-01-24 v2
Abstract
The standard assumption in reinforcement learning (RL) is that agents observe feedback for their actions immediately. However, in practice feedback is often observed in delay. This paper studies online learning in episodic Markov decision process (MDP) with unknown transitions, adversarially changing costs, and unrestricted delayed bandit feedback. More precisely, the feedback for the agent in episode is revealed only in the end of episode , where the delay can be changing over episodes and chosen by an oblivious adversary. We present the first algorithms that achieve near-optimal regret, where is the number of episodes and is the total delay, significantly improving upon the best known regret bound of .
Keywords
Cite
@article{arxiv.2201.13172,
title = {Near-Optimal Regret for Adversarial MDP with Delayed Bandit Feedback},
author = {Tiancheng Jin and Tal Lancewicki and Haipeng Luo and Yishay Mansour and Aviv Rosenberg},
journal= {arXiv preprint arXiv:2201.13172},
year = {2023}
}
Comments
NeurIPS 2022