English

Improved Regret in Stochastic Decision-Theoretic Online Learning under Differential Privacy

Machine Learning 2025-06-19 v2 Cryptography and Security Data Structures and Algorithms

Abstract

Hu and Mehta (2024) posed an open problem: what is the optimal instance-dependent rate for the stochastic decision-theoretic online learning (with KK actions and TT rounds) under ε\varepsilon-differential privacy? Before, the best known upper bound and lower bound are O(logKΔmin+logKlogTε)O\left(\frac{\log K}{\Delta_{\min}} + \frac{\log K\log T}{\varepsilon}\right) and Ω(logKΔmin+logKε)\Omega\left(\frac{\log K}{\Delta_{\min}} + \frac{\log K}{\varepsilon}\right) (where Δmin\Delta_{\min} is the gap between the optimal and the second actions). In this paper, we partially address this open problem by having two new results. First, we provide an improved upper bound for this problem O(logKΔmin+log2Kε)O\left(\frac{\log K}{\Delta_{\min}} + \frac{\log^2K}{\varepsilon}\right), which is TT-independent and only has a log dependency in KK. Second, to further understand the gap, we introduce the \textit{deterministic setting}, a weaker setting of this open problem, where the received loss vector is deterministic. At this weaker setting, a direct application of the analysis and algorithms from the original setting still leads to an extra log factor. We conduct a novel analysis which proves upper and lower bounds that match at Θ(logKε)\Theta(\frac{\log K}{\varepsilon}).

Keywords

Cite

@article{arxiv.2502.10997,
  title  = {Improved Regret in Stochastic Decision-Theoretic Online Learning under Differential Privacy},
  author = {Ruihan Wu and Yu-Xiang Wang},
  journal= {arXiv preprint arXiv:2502.10997},
  year   = {2025}
}
R2 v1 2026-06-28T21:45:47.069Z