English

Stochastic $k$-Submodular Bandits with Full Bandit Feedback

Machine Learning 2024-12-17 v1 Data Structures and Algorithms Machine Learning

Abstract

In this paper, we present the first sublinear α\alpha-regret bounds for online kk-submodular optimization problems with full-bandit feedback, where α\alpha is a corresponding offline approximation ratio. Specifically, we propose online algorithms for multiple kk-submodular stochastic combinatorial multi-armed bandit problems, including (i) monotone functions and individual size constraints, (ii) monotone functions with matroid constraints, (iii) non-monotone functions with matroid constraints, (iv) non-monotone functions without constraints, and (v) monotone functions without constraints. We transform approximation algorithms for offline kk-submodular maximization problems into online algorithms through the offline-to-online framework proposed by Nie et al. (2023a). A key contribution of our work is analyzing the robustness of the offline algorithms.

Keywords

Cite

@article{arxiv.2412.10682,
  title  = {Stochastic $k$-Submodular Bandits with Full Bandit Feedback},
  author = {Guanyu Nie and Vaneet Aggarwal and Christopher John Quinn},
  journal= {arXiv preprint arXiv:2412.10682},
  year   = {2024}
}

Comments

26 pages, 1 figure

R2 v1 2026-06-28T20:35:00.372Z