Improved Regret Bounds for Bandits with Expert Advice
Machine Learning
2024-06-25 v1 Machine Learning
Abstract
In this research note, we revisit the bandits with expert advice problem. Under a restricted feedback model, we prove a lower bound of order for the worst-case regret, where is the number of actions, the number of experts, and the time horizon. This matches a previously known upper bound of the same order and improves upon the best available lower bound of . For the standard feedback model, we prove a new instance-based upper bound that depends on the agreement between the experts and provides a logarithmic improvement compared to prior results.
Keywords
Cite
@article{arxiv.2406.16802,
title = {Improved Regret Bounds for Bandits with Expert Advice},
author = {Nicolò Cesa-Bianchi and Khaled Eldowa and Emmanuel Esposito and Julia Olkhovskaya},
journal= {arXiv preprint arXiv:2406.16802},
year = {2024}
}