具有对抗性奖励和赌博机反馈的确定性马尔可夫决策过程
计算机科学与博弈论
2012-10-19 v1 机器学习
摘要
我们考虑一个马尔可夫决策过程,具有确定性状态转移动力学、对抗性生成的奖励(每轮任意变化)以及赌博机反馈模型(决策者只观察到它获得的奖励)。在这种设置下,我们提出了一种新颖且高效的在线决策算法,名为MarcoPolo。在转移动力学结构的温和假设下,我们证明MarcoPolo相对于事后最佳确定性策略的遗憾为O(T^(3/4)sqrt(log(T)))。具体来说,我们的分析不依赖于严格的单链假设,而这一假设主导了该主题的大部分先前工作。
引用
@article{arxiv.1210.4843,
title = {Deterministic MDPs with Adversarial Rewards and Bandit Feedback},
author = {Raman Arora and Ofer Dekel and Ambuj Tewari},
journal= {arXiv preprint arXiv:1210.4843},
year = {2012}
}
备注
Appears in Proceedings of the Twenty-Eighth Conference on Uncertainty in Artificial Intelligence (UAI2012)