A Diffusion Analysis of Policy Gradient for Stochastic Bandits
Machine Learning
2026-03-12 v1 Artificial Intelligence
Machine Learning
Statistics Theory
Statistics Theory
Abstract
We study a continuous-time diffusion approximation of policy gradient for -armed stochastic bandits. We prove that with a learning rate the regret is where is the horizon and the minimum gap. Moreover, we construct an instance with only logarithmically many arms for which the regret is linear unless .
Keywords
Cite
@article{arxiv.2603.10219,
title = {A Diffusion Analysis of Policy Gradient for Stochastic Bandits},
author = {Tor Lattimore},
journal= {arXiv preprint arXiv:2603.10219},
year = {2026}
}
Comments
17 pages