English

Stochastic Bandits with ReLU Neural Networks

Machine Learning 2024-05-14 v1 Data Structures and Algorithms Machine Learning

Abstract

We study the stochastic bandit problem with ReLU neural network structure. We show that a O~(T)\tilde{O}(\sqrt{T}) regret guarantee is achievable by considering bandits with one-layer ReLU neural networks; to the best of our knowledge, our work is the first to achieve such a guarantee. In this specific setting, we propose an OFU-ReLU algorithm that can achieve this upper bound. The algorithm first explores randomly until it reaches a linear regime, and then implements a UCB-type linear bandit algorithm to balance exploration and exploitation. Our key insight is that we can exploit the piecewise linear structure of ReLU activations and convert the problem into a linear bandit in a transformed feature space, once we learn the parameters of ReLU relatively accurately during the exploration stage. To remove dependence on model parameters, we design an OFU-ReLU+ algorithm based on a batching strategy, which can provide the same theoretical guarantee.

Keywords

Cite

@article{arxiv.2405.07331,
  title  = {Stochastic Bandits with ReLU Neural Networks},
  author = {Kan Xu and Hamsa Bastani and Surbhi Goel and Osbert Bastani},
  journal= {arXiv preprint arXiv:2405.07331},
  year   = {2024}
}