English

PILLAR: How to make semi-private learning more effective

Machine Learning 2023-06-08 v1 Artificial Intelligence Cryptography and Security Machine Learning

Abstract

In Semi-Supervised Semi-Private (SP) learning, the learner has access to both public unlabelled and private labelled data. We propose a computationally efficient algorithm that, under mild assumptions on the data, provably achieves significantly lower private labelled sample complexity and can be efficiently run on real-world datasets. For this purpose, we leverage the features extracted by networks pre-trained on public (labelled or unlabelled) data, whose distribution can significantly differ from the one on which SP learning is performed. To validate its empirical effectiveness, we propose a wide variety of experiments under tight privacy constraints (ϵ=0.1\epsilon = 0.1) and with a focus on low-data regimes. In all of these settings, our algorithm exhibits significantly improved performance over available baselines that use similar amounts of public data.

Keywords

Cite

@article{arxiv.2306.03962,
  title  = {PILLAR: How to make semi-private learning more effective},
  author = {Francesco Pinto and Yaxi Hu and Fanny Yang and Amartya Sanyal},
  journal= {arXiv preprint arXiv:2306.03962},
  year   = {2023}
}
R2 v1 2026-06-28T10:58:11.346Z