English

Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits

Machine Learning 2024-11-18 v4 Statistics Theory Machine Learning Statistics Theory

Abstract

We address the problem of stochastic combinatorial semi-bandits, where a player selects among P actions from the power set of a set containing d base items. Adaptivity to the problem's structure is essential in order to obtain optimal regret upper bounds. As estimating the coefficients of a covariance matrix can be manageable in practice, leveraging them should improve the regret. We design "optimistic" covariance-adaptive algorithms relying on online estimations of the covariance structure, called OLS-UCB-C and COS-V (only the variances for the latter). They both yields improved gap-free regret. Although COS-V can be slightly suboptimal, it improves on computational complexity by taking inspiration from ThompsonSampling approaches. It is the first sampling-based algorithm satisfying a T^1/2 gap-free regret (up to poly-logs). We also show that in some cases, our approach efficiently leverages the semi-bandit feedback and outperforms bandit feedback approaches, not only in exponential regimes where P >> d but also when P <= d, which is not covered by existing analyses.

Keywords

Cite

@article{arxiv.2402.15171,
  title  = {Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits},
  author = {Julien Zhou and Pierre Gaillard and Thibaud Rahier and Houssam Zenati and Julyan Arbel},
  journal= {arXiv preprint arXiv:2402.15171},
  year   = {2024}
}
R2 v1 2026-06-28T14:58:06.819Z