English

Bandit Multi-linear DR-Submodular Maximization and Its Applications on Adversarial Submodular Bandits

Machine Learning 2023-05-23 v1 Artificial Intelligence

Abstract

We investigate the online bandit learning of the monotone multi-linear DR-submodular functions, designing the algorithm BanditMLSM\mathtt{BanditMLSM} that attains O(T2/3logT)O(T^{2/3}\log T) of (11/e)(1-1/e)-regret. Then we reduce submodular bandit with partition matroid constraint and bandit sequential monotone maximization to the online bandit learning of the monotone multi-linear DR-submodular functions, attaining O(T2/3logT)O(T^{2/3}\log T) of (11/e)(1-1/e)-regret in both problems, which improve the existing results. To the best of our knowledge, we are the first to give a sublinear regret algorithm for the submodular bandit with partition matroid constraint. A special case of this problem is studied by Streeter et al.(2009). They prove a O(T4/5)O(T^{4/5}) (11/e)(1-1/e)-regret upper bound. For the bandit sequential submodular maximization, the existing work proves an O(T2/3)O(T^{2/3}) regret with a suboptimal 1/21/2 approximation ratio (Niazadeh et al. 2021).

Keywords

Cite

@article{arxiv.2305.12402,
  title  = {Bandit Multi-linear DR-Submodular Maximization and Its Applications on Adversarial Submodular Bandits},
  author = {Zongqi Wan and Jialin Zhang and Wei Chen and Xiaoming Sun and Zhijie Zhang},
  journal= {arXiv preprint arXiv:2305.12402},
  year   = {2023}
}

Comments

Accepted by ICML 2023

R2 v1 2026-06-28T10:40:25.160Z