English

Multi-Armed Bandits and Quantum Channel Oracles

Quantum Physics 2025-03-26 v4 Machine Learning

Abstract

Multi-armed bandits are one of the theoretical pillars of reinforcement learning. Recently, the investigation of quantum algorithms for multi-armed bandit problems was started, and it was found that a quadratic speed-up (in query complexity) is possible when the arms and the randomness of the rewards of the arms can be queried in superposition. Here we introduce further bandit models where we only have limited access to the randomness of the rewards, but we can still query the arms in superposition. We show that then the query complexity is the same as for classical algorithms. This generalizes the prior result that no speed-up is possible for unstructured search when the oracle has positive failure probability.

Keywords

Cite

@article{arxiv.2301.08544,
  title  = {Multi-Armed Bandits and Quantum Channel Oracles},
  author = {Simon Buchholz and Jonas M. Kübler and Bernhard Schölkopf},
  journal= {arXiv preprint arXiv:2301.08544},
  year   = {2025}
}

Comments

50 pages, accepted in Quantum

R2 v1 2026-06-28T08:16:09.439Z