English

Beyond $\tilde{O}(\sqrt{T})$ Constraint Violation for Online Convex Optimization with Adversarial Constraints

Machine Learning 2025-11-17 v2 Optimization and Control Machine Learning

Abstract

We study Online Convex Optimization with adversarial constraints (COCO). At each round a learner selects an action from a convex decision set and then an adversary reveals a convex cost and a convex constraint function. The goal of the learner is to select a sequence of actions to minimize both regret and the cumulative constraint violation (CCV) over a horizon of length TT. The best-known policy for this problem achieves O(T)O(\sqrt{T}) regret and O~(T)\tilde{O}(\sqrt{T}) CCV. In this paper, we improve this by trading off regret to achieve substantially smaller CCV. This trade-off is especially important in safety-critical applications, where satisfying the safety constraints is non-negotiable. Specifically, for any bounded convex cost and constraint functions, we propose an online policy that achieves O~(dT+Tβ)\tilde{O}(\sqrt{dT}+ T^\beta) regret and O~(dT1β)\tilde{O}(dT^{1-\beta}) CCV, where dd is the dimension of the decision set and β[0,1]\beta \in [0,1] is a tunable parameter. We begin with a special case, called the Constrained Expert\textsf{Constrained Expert} problem, where the decision set is a probability simplex and the cost and constraint functions are linear. Leveraging a new adaptive small-loss regret bound, we propose a computationally efficient policy for the Constrained Expert\textsf{Constrained Expert} problem, that attains O(TlnN+Tβ)O(\sqrt{T\ln N}+T^{\beta}) regret and O~(T1βlnN)\tilde{O}(T^{1-\beta} \ln N) CCV for NN number of experts. The original problem is then reduced to the Constrained Expert\textsf{Constrained Expert} problem via a covering argument. Finally, with an additional MM-smoothness assumption, we propose a computationally efficient first-order policy attaining O(MT+Tβ)O(\sqrt{MT}+T^{\beta}) regret and O~(MT1β)\tilde{O}(MT^{1-\beta}) CCV.

Keywords

Cite

@article{arxiv.2505.06709,
  title  = {Beyond $\tilde{O}(\sqrt{T})$ Constraint Violation for Online Convex Optimization with Adversarial Constraints},
  author = {Abhishek Sinha and Rahul Vaze},
  journal= {arXiv preprint arXiv:2505.06709},
  year   = {2025}
}

Comments

To appear in NeurIPS 2025