English

Online Convex Optimization with Stochastic Constraints: Zero Constraint Violation and Bandit Feedback

Optimization and Control 2023-07-17 v2 Machine Learning

Abstract

This paper studies online convex optimization with stochastic constraints. We propose a variant of the drift-plus-penalty algorithm that guarantees O(T)O(\sqrt{T}) expected regret and zero constraint violation, after a fixed number of iterations, which improves the vanilla drift-plus-penalty method with O(T)O(\sqrt{T}) constraint violation. Our algorithm is oblivious to the length of the time horizon TT, in contrast to the vanilla drift-plus-penalty method. This is based on our novel drift lemma that provides time-varying bounds on the virtual queue drift and, as a result, leads to time-varying bounds on the expected virtual queue length. Moreover, we extend our framework to stochastic-constrained online convex optimization under two-point bandit feedback. We show that by adapting our algorithmic framework to the bandit feedback setting, we may still achieve O(T)O(\sqrt{T}) expected regret and zero constraint violation, improving upon the previous work for the case of identical constraint functions. Numerical results demonstrate our theoretical results.

Keywords

Cite

@article{arxiv.2301.11267,
  title  = {Online Convex Optimization with Stochastic Constraints: Zero Constraint Violation and Bandit Feedback},
  author = {Yeongjong Kim and Dabeen Lee},
  journal= {arXiv preprint arXiv:2301.11267},
  year   = {2023}
}

Comments

We found a paper that has already obtained the results of the submission

R2 v1 2026-06-28T08:22:00.899Z