English

A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits

Machine Learning 2026-03-30 v1

Abstract

We adapt the analysis of policy gradient for continuous time kk-armed stochastic bandits by Lattimore (2026) to the standard discrete time setup. As in continuous time, we prove that with learning rate η=O(Δmin2/(Δmaxlog(n)))\eta = O(\Delta_{\min}^2/(\Delta_{\max} \log(n))) the regret is O(klog(k)log(n)/η)O(k \log(k) \log(n) / \eta) where nn is the horizon and Δmin\Delta_{\min} and Δmax\Delta_{\max} are the minimum and maximum gaps.

Keywords

Cite

@article{arxiv.2603.26547,
  title  = {A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits},
  author = {Tor Lattimore},
  journal= {arXiv preprint arXiv:2603.26547},
  year   = {2026}
}

Comments

6 pages