English

Online optimization and regret guarantees for non-additive long-term constraints

Machine Learning 2016-06-09 v2 Machine Learning Optimization and Control Statistics Theory Statistics Theory

Abstract

We consider online optimization in the 1-lookahead setting, where the objective does not decompose additively over the rounds of the online game. The resulting formulation enables us to deal with non-stationary and/or long-term constraints , which arise, for example, in online display advertising problems. We propose an on-line primal-dual algorithm for which we obtain dynamic cumulative regret guarantees. They depend on the convexity and the smoothness of the non-additive penalty, as well as terms capturing the smoothness with which the residuals of the non-stationary and long-term constraints vary over the rounds. We conduct experiments on synthetic data to illustrate the benefits of the non-additive penalty and show vanishing regret convergence on live traffic data collected by a display advertising platform in production.

Keywords

Cite

@article{arxiv.1602.05394,
  title  = {Online optimization and regret guarantees for non-additive long-term constraints},
  author = {Rodolphe Jenatton and Jim Huang and Dominik Csiba and Cedric Archambeau},
  journal= {arXiv preprint arXiv:1602.05394},
  year   = {2016}
}
R2 v1 2026-06-22T12:52:08.888Z