English

Identifying the Best Transition Law

Machine Learning 2025-02-19 v1 Artificial Intelligence

Abstract

Motivated by recursive learning in Markov Decision Processes, this paper studies best-arm identification in bandit problems where each arm's reward is drawn from a multinomial distribution with a known support. We compare the performance { reached by strategies including notably LUCB without and with use of this knowledge. } In the first case, we use classical non-parametric approaches for the confidence intervals. In the second case, where a probability distribution is to be estimated, we first use classical deviation bounds (Hoeffding and Bernstein) on each dimension independently, and then the Empirical Likelihood method (EL-LUCB) on the joint probability vector. The effectiveness of these methods is demonstrated through simulations on scenarios with varying levels of structural complexity.

Keywords

Cite

@article{arxiv.2502.12227,
  title  = {Identifying the Best Transition Law},
  author = {Mehrasa Ahmadipour and élise Crepon and Aurélien Garivier},
  journal= {arXiv preprint arXiv:2502.12227},
  year   = {2025}
}
R2 v1 2026-06-28T21:47:48.394Z