English

Optimal Gap-Dependent Regret for Private Stochastic Decision-Theoretic Online Learning

Machine Learning 2026-05-29 v1 Machine Learning

Abstract

We study stochastic decision-theoretic online learning with full information and event-level pure differential privacy. A COLT open problem of Hu and Mehta asks to determine the optimal gap-dependent regret rate for stochastic decision-theoretic online learning under pure event-level differential privacy. For KK actions, losses in [0,1][0,1], and a unique best action separated from the second-best action by gap Δmin\Delta_{\min}, the known lower bound is of order logKmin{Δmin,ε}, \frac{\log K}{\min\{\Delta_{\min},\varepsilon\}}, or equivalently, up to universal constants, of order logKΔmin+logKε. \frac{\log K}{\Delta_{\min}}+\frac{\log K}{\varepsilon}. We give a horizon-free pure-DP algorithm and prove the explicit regret bound RegT1000(logKΔmin+logKε) \operatorname{Reg}_T \le 1000 \cdot \left(\frac{\log K}{\Delta_{\min}}+\frac{\log K}{\varepsilon}\right) for every horizon TT. The numerical constant is not optimized. The algorithm partitions time into blocks of exponentially increasing size, plays a single action throughout each block, and chooses the next action by an exponential mechanism applied to a data-independent random prefix of the previous block. The random prefix converts block regret into a sum, over all prefix lengths, of softmax selection errors. A single entropy-potential argument controls all privacy-dominated large-gap actions at cost logK/ε\log K/\varepsilon.

Keywords

Cite

@article{arxiv.2605.29148,
  title  = {Optimal Gap-Dependent Regret for Private Stochastic Decision-Theoretic Online Learning},
  author = {Tommaso Cesari and Roberto Colomboni},
  journal= {arXiv preprint arXiv:2605.29148},
  year   = {2026}
}
R2 v1 2026-07-22T07:38:21.471Z