English

Minimax Regret for Bandit Convex Optimisation of Ridge Functions

Machine Learning 2021-06-08 v2 Optimization and Control

Abstract

We analyse adversarial bandit convex optimisation with an adversary that is restricted to playing functions of the form ft(x)=gt(x,θ)f_t(x) = g_t(\langle x, \theta\rangle) for convex gt:RRg_t : \mathbb R \to \mathbb R and unknown θRd\theta \in \mathbb R^d that is homogeneous over time. We provide a short information-theoretic proof that the minimax regret is at most O(dnlog(ndiam(K)))O(d \sqrt{n} \log(n \operatorname{diam}(\mathcal K))) where nn is the number of interactions, dd the dimension and diam(K)\operatorname{diam}(\mathcal K) is the diameter of the constraint set.

Keywords

Cite

@article{arxiv.2106.00444,
  title  = {Minimax Regret for Bandit Convex Optimisation of Ridge Functions},
  author = {Tor Lattimore},
  journal= {arXiv preprint arXiv:2106.00444},
  year   = {2021}
}

Comments

Correcting an (instructive) error that leads to a weaker result