English

Generalised Entropy MDPs and Minimax Regret

Machine Learning 2014-12-11 v1 Machine Learning

Abstract

Bayesian methods suffer from the problem of how to specify prior beliefs. One interesting idea is to consider worst-case priors. This requires solving a stochastic zero-sum game. In this paper, we extend well-known results from bandit theory in order to discover minimax-Bayes policies and discuss when they are practical.

Cite

@article{arxiv.1412.3276,
  title  = {Generalised Entropy MDPs and Minimax Regret},
  author = {Emmanouil G. Androulakis and Christos Dimitrakakis},
  journal= {arXiv preprint arXiv:1412.3276},
  year   = {2014}
}

Comments

7 pages, NIPS workshop "From bad models to good policies"

R2 v1 2026-06-22T07:26:22.328Z