English

A Diffusion Analysis of Policy Gradient for Stochastic Bandits

Machine Learning 2026-03-12 v1 Artificial Intelligence Machine Learning Statistics Theory Statistics Theory

Abstract

We study a continuous-time diffusion approximation of policy gradient for kk-armed stochastic bandits. We prove that with a learning rate η=O(Δ2/log(n))\eta = O(\Delta^2/\log(n)) the regret is O(klog(k)log(n)/η)O(k \log(k) \log(n) / \eta) where nn is the horizon and Δ\Delta the minimum gap. Moreover, we construct an instance with only logarithmically many arms for which the regret is linear unless η=O(Δ2)\eta = O(\Delta^2).

Keywords

Cite

@article{arxiv.2603.10219,
  title  = {A Diffusion Analysis of Policy Gradient for Stochastic Bandits},
  author = {Tor Lattimore},
  journal= {arXiv preprint arXiv:2603.10219},
  year   = {2026}
}

Comments

17 pages

R2 v1 2026-07-01T11:13:51.466Z