Online Learning for Non-monotone Submodular Maximization: From Full Information to Bandit Feedback
Abstract
In this paper, we revisit the online non-monotone continuous DR-submodular maximization problem over a down-closed convex set, which finds wide real-world applications in the domain of machine learning, economics, and operations research. At first, we present the Meta-MFW algorithm achieving a -regret of at the cost of stochastic gradient evaluations per round. As far as we know, Meta-MFW is the first algorithm to obtain -regret of for the online non-monotone continuous DR-submodular maximization problem over a down-closed convex set. Furthermore, in sharp contrast with ODC algorithm \citep{thang2021online}, Meta-MFW relies on the simple online linear oracle without discretization, lifting, or rounding operations. Considering the practical restrictions, we then propose the Mono-MFW algorithm, which reduces the per-function stochastic gradient evaluations from to 1 and achieves a -regret bound of . Next, we extend Mono-MFW to the bandit setting and propose the Bandit-MFW algorithm which attains a -regret bound of . To the best of our knowledge, Mono-MFW and Bandit-MFW are the first sublinear-regret algorithms to explore the one-shot and bandit setting for online non-monotone continuous DR-submodular maximization problem over a down-closed convex set, respectively. Finally, we conduct numerical experiments on both synthetic and real-world datasets to verify the effectiveness of our methods.
Keywords
Cite
@article{arxiv.2208.07632,
title = {Online Learning for Non-monotone Submodular Maximization: From Full Information to Bandit Feedback},
author = {Qixin Zhang and Zengde Deng and Zaiyi Chen and Kuangqi Zhou and Haoyuan Hu and Yu Yang},
journal= {arXiv preprint arXiv:2208.07632},
year = {2022}
}
Comments
31 pages, 6 figures, 3 tables