English

On Private and Robust Bandits

Machine Learning 2023-03-07 v2 Artificial Intelligence Cryptography and Security Machine Learning

Abstract

We study private and robust multi-armed bandits (MABs), where the agent receives Huber's contaminated heavy-tailed rewards and meanwhile needs to ensure differential privacy. We first present its minimax lower bound, characterizing the information-theoretic limit of regret with respect to privacy budget, contamination level and heavy-tailedness. Then, we propose a meta-algorithm that builds on a private and robust mean estimation sub-routine \texttt{PRM} that essentially relies on reward truncation and the Laplace mechanism only. For two different heavy-tailed settings, we give specific schemes of \texttt{PRM}, which enable us to achieve nearly-optimal regret. As by-products of our main results, we also give the first minimax lower bound for private heavy-tailed MABs (i.e., without contamination). Moreover, our two proposed truncation-based \texttt{PRM} achieve the optimal trade-off between estimation accuracy, privacy and robustness. Finally, we support our theoretical results with experimental studies.

Keywords

Cite

@article{arxiv.2302.02526,
  title  = {On Private and Robust Bandits},
  author = {Yulian Wu and Xingyu Zhou and Youming Tao and Di Wang},
  journal= {arXiv preprint arXiv:2302.02526},
  year   = {2023}
}
R2 v1 2026-06-28T08:32:35.542Z