English
Related papers

Related papers: A False Sense of Privacy: Evaluating Textual Data …

200 papers

Language models are widely deployed to provide automatic text completion services in user products. However, recent research has revealed that language models (especially large ones) bear considerable risk of memorizing private training…

Computation and Language · Computer Science 2022-12-19 C. M. Downey , Wei Dai , Huseyin A. Inan , Kim Laine , Saurabh Naik , Tomasz Religa

Preserving privacy of continuous and/or high-dimensional data such as images, videos and audios, can be challenging with syntactic anonymization methods which are designed for discrete attributes. Differential privacy, which provides a more…

Machine Learning · Computer Science 2017-12-04 Jihun Hamm

Differential Privacy (DP) for text matured from disjointed word-level substitutions to contiguous sentence-level rewriting by leveraging the generative capacity of language models. While this form of text privatization is best suited for…

Computation and Language · Computer Science 2026-04-30 Stefan Arnold

As the issues of privacy and trust are receiving increasing attention within the research community, various attempts have been made to anonymize textual data. A significant subset of these approaches incorporate differentially private…

Cryptography and Security · Computer Science 2022-05-05 Justus Mattern , Benjamin Weggenmann , Florian Kerschbaum

Text de-identification techniques are often used to mask personally identifiable information (PII) from documents. Their ability to conceal the identity of the individuals mentioned in a text is, however, hard to measure. Recent work has…

Computation and Language · Computer Science 2025-10-13 Lucas Georges Gabriel Charpentier , Pierre Lison

Browsing-enabled LLM assistants can fetch webpages and answer contact-seeking queries, creating a practical channel for scraping contact-style personally identifiable information (PII) from public pages. Many prior defenses are deployed at…

Cryptography and Security · Computer Science 2026-05-06 Mingshuo Liu , Yiwei Zha , Min Chen

Unprecedented data collection and sharing have exacerbated privacy concerns and led to increasing interest in privacy-preserving tools that remove sensitive attributes from images while maintaining useful information for other tasks.…

Computer Vision and Pattern Recognition · Computer Science 2020-09-22 Kang Liu , Benjamin Tan , Siddharth Garg

Personal data collected at scale promises to improve decision-making and accelerate innovation. However, sharing and using such data raises serious privacy concerns. A promising solution is to produce synthetic data, artificial records to…

This paper presents the first large-scale empirical study of commercial personally identifiable information (PII) removal systems -- commercial services that claim to improve privacy by automating the removal of PII from data broker's…

Cryptography and Security · Computer Science 2025-07-15 Jiahui He , Pete Snyder , Hamed Haddadi , Fabián E. Bustamante , Gareth Tyson

Removing Personally Identifiable Information (PII) from clinical notes in Electronic Health Records (EHRs) is essential for research and AI development. While Large Language Models (LLMs) are powerful, their high computational costs and the…

Computation and Language · Computer Science 2025-10-23 Prakrithi Shivaprakash , Lekhansh Shukla , Animesh Mukherjee , Prabhat Chand , Pratima Murthy

The field of text privatization often leverages the notion of $\textit{Differential Privacy}$ (DP) to provide formal guarantees in the rewriting or obfuscation of sensitive textual data. A common and nearly ubiquitous form of DP application…

Computation and Language · Computer Science 2025-02-03 Stephen Meisenbacher , Maulik Chevli , Florian Matthes

Protecting Personal Identifiable Information (PII) in text data is crucial for privacy, but current PII generalization methods face challenges such as uneven data distributions and limited context awareness. To address these issues, we…

Computation and Language · Computer Science 2024-07-04 Kailin Zhang , Xinying Qiu

The generation of privacy-preserving synthetic datasets is a promising avenue for overcoming data scarcity in medical AI research. Post-hoc privacy filtering techniques, designed to remove samples containing personally identifiable…

Machine Learning · Computer Science 2025-10-03 Adil Koeken , Alexander Ziller , Moritz Knolle , Daniel Rueckert

With the growing use of camera devices, the industry has many image datasets that provide more opportunities for collaboration between the machine learning community and industry. However, the sensitive information in the datasets…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Jia-Wei Chen , Li-Ju Chen , Chia-Mu Yu , Chun-Shien Lu

To prove that a dataset is sufficiently anonymized, many privacy policies suggest that a re-identification risk assessment be performed, but do not provide a precise methodology for doing so, leaving the industry alone with the problem.…

Cryptography and Security · Computer Science 2025-01-22 Louis-Philippe Sondeck , Maryline Laurent

Anonymization technique has been extensively studied and widely applied for privacy-preserving data publishing. In most previous approaches, a microdata table consists of three categories of attribute: explicit-identifier, quasi-identifier…

Cryptography and Security · Computer Science 2020-08-26 Boyu Li , Kun He , Geng Sun

As Large Language Models (LLMs) achieve remarkable success across a wide range of applications, such as chatbots and code copilots, concerns surrounding the generation of harmful content have come increasingly into focus. Despite…

Computation and Language · Computer Science 2025-09-30 Wenjie Fu , Huandong Wang , Junyao Gao , Guoan Wan , Tao Jiang

This paper proposes and compares measures of identity and attribute disclosure risk for synthetic data. Data custodians can use the methods proposed here to inform the decision as to whether to release synthetic versions of confidential…

Applications · Statistics 2025-05-19 Gillian M Raab

This systematic literature review investigates perceptions, concerns, and expectations of young digital citizens regarding privacy in artificial intelligence (AI) systems, focusing on social media platforms, educational technology, gaming…

Computers and Society · Computer Science 2025-12-16 Ajay Kumar Shrestha , Ankur Barthwal , Molly Campbell , Austin Shouli , Saad Syed , Sandhya Joshi , Julita Vassileva

Data anonymization is an approach to privacy-preserving data release aimed at preventing participants reidentification, and it is an important alternative to differential privacy in applications that cannot tolerate noisy data. Existing…

Data Structures and Algorithms · Computer Science 2022-01-31 Gecia Bravo-Hermsdorff , Robert Busa-Fekete , Lee M. Gunderson , Andrés Munõz Medina , Umar Syed