English

Thompson Sampling under Bernoulli Rewards with Local Differential Privacy

Machine Learning 2023-07-04 v1 Cryptography and Security

Abstract

This paper investigates the problem of regret minimization for multi-armed bandit (MAB) problems with local differential privacy (LDP) guarantee. Given a fixed privacy budget ϵ\epsilon, we consider three privatizing mechanisms under Bernoulli scenario: linear, quadratic and exponential mechanisms. Under each mechanism, we derive stochastic regret bound for Thompson Sampling algorithm. Finally, we simulate to illustrate the convergence of different mechanisms under different privacy budgets.

Keywords

Cite

@article{arxiv.2307.00863,
  title  = {Thompson Sampling under Bernoulli Rewards with Local Differential Privacy},
  author = {Bo Jiang and Tianchi Zhao and Ming Li},
  journal= {arXiv preprint arXiv:2307.00863},
  year   = {2023}
}

Comments

Accepted by ICML 22 workshop

R2 v1 2026-06-28T11:20:32.197Z