English

Differential Privacy for Multi-armed Bandits: What Is It and What Is Its Cost?

Machine Learning 2020-06-25 v2 Machine Learning

Abstract

Based on differential privacy (DP) framework, we introduce and unify privacy definitions for the multi-armed bandit algorithms. We represent the framework with a unified graphical model and use it to connect privacy definitions. We derive and contrast lower bounds on the regret of bandit algorithms satisfying these definitions. We leverage a unified proving technique to achieve all the lower bounds. We show that for all of them, the learner's regret is increased by a multiplicative factor dependent on the privacy level ϵ\epsilon. We observe that the dependency is weaker when we do not require local differential privacy for the rewards.

Keywords

Cite

@article{arxiv.1905.12298,
  title  = {Differential Privacy for Multi-armed Bandits: What Is It and What Is Its Cost?},
  author = {Debabrota Basu and Christos Dimitrakakis and Aristide Tossou},
  journal= {arXiv preprint arXiv:1905.12298},
  year   = {2020}
}

Comments

27 pages, 1 figure, 2 tables, 14 theorems