Efficient Online-Bandit Strategies for Minimax Learning Problems
Abstract
Several learning problems involve solving min-max problems, e.g., empirical distributional robust learning or learning with non-standard aggregated losses. More specifically, these problems are convex-linear problems where the minimization is carried out over the model parameters and the maximization over the empirical distribution of the training set indexes, where is the simplex or a subset of it. To design efficient methods, we let an online learning algorithm play against a (combinatorial) bandit algorithm. We argue that the efficiency of such approaches critically depends on the structure of and propose two properties of that facilitate designing efficient algorithms. We focus on a specific family of sets encompassing various learning applications and provide high-probability convergence guarantees to the minimax values.
Cite
@article{arxiv.2105.13939,
title = {Efficient Online-Bandit Strategies for Minimax Learning Problems},
author = {Christophe Roux and Elias Wirth and Sebastian Pokutta and Thomas Kerdreux},
journal= {arXiv preprint arXiv:2105.13939},
year = {2021}
}