English

Differential Privacy for Text Analytics via Natural Text Sanitization

Computation and Language 2021-06-03 v1 Cryptography and Security

Abstract

Texts convey sophisticated knowledge. However, texts also convey sensitive information. Despite the success of general-purpose language models and domain-specific mechanisms with differential privacy (DP), existing text sanitization mechanisms still provide low utility, as cursed by the high-dimensional text representation. The companion issue of utilizing sanitized texts for downstream analytics is also under-explored. This paper takes a direct approach to text sanitization. Our insight is to consider both sensitivity and similarity via our new local DP notion. The sanitized texts also contribute to our sanitization-aware pretraining and fine-tuning, enabling privacy-preserving natural language processing over the BERT language model with promising utility. Surprisingly, the high utility does not boost up the success rate of inference attacks.

Keywords

Cite

@article{arxiv.2106.01221,
  title  = {Differential Privacy for Text Analytics via Natural Text Sanitization},
  author = {Xiang Yue and Minxin Du and Tianhao Wang and Yaliang Li and Huan Sun and Sherman S. M. Chow},
  journal= {arXiv preprint arXiv:2106.01221},
  year   = {2021}
}

Comments

ACL-ICJNLP'21 Findings; The first two authors contributed equally

R2 v1 2026-06-24T02:45:17.195Z