English
Related papers

Related papers: PIIvot: A Lightweight NLP Anonymization Framework …

200 papers

This paper studies a novel privacy-preserving anonymization problem for pedestrian images, which preserves personal identity information (PII) for authorized models and prevents PII from being recognized by third parties. Conventional…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Junwu Zhang , Mang Ye , Yao Yang

Redacting Personally Identifiable Information (PII) from unstructured text is critical for ensuring data privacy in regulated domains. While earlier approaches have relied on rule-based systems and domain-specific Named Entity Recognition…

Cryptography and Security · Computer Science 2025-08-08 Leon Garza , Anantaa Kotal , Aritran Piplai , Lavanya Elluri , Prajit Das , Aman Chadha

Prompt serves as a crucial link in interacting with large language models (LLMs), widely impacting the accuracy and interpretability of model outputs. However, acquiring accurate and high-quality responses necessitates precise prompts,…

Cryptography and Security · Computer Science 2024-08-20 Xiongtao Sun , Gan Liu , Zhipeng He , Hui Li , Xiaoguang Li

Concerns regarding Large Language Models (LLMs) to memorize and disclose private information, particularly Personally Identifiable Information (PII), become prominent within the community. Many efforts have been made to mitigate the privacy…

Machine Learning · Computer Science 2024-05-21 Ruizhe Chen , Tianxiang Hu , Yang Feng , Zuozhu Liu

De-identification of data used for automatic speech recognition modeling is a critical component in protecting privacy, especially in the medical domain. However, simply removing all personally identifiable information (PII) from end-to-end…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-13 Martin Flechl , Shou-Chun Yin , Junho Park , Peter Skala

As Large Language Models (LLMs) gain wider adoption, ensuring their reliable handling of Personally Identifiable Information (PII) across diverse regulatory contexts has become essential. This work introduces a scalable multilingual data…

Computation and Language · Computer Science 2025-10-13 Bharti Meena , Joanna Skubisz , Harshit Rajgarhia , Nand Dave , Kiran Ganesh , Shivali Dalmia , Abhishek Mukherji , Vasudevan Sundarababu

Current text anonymization evaluation relies on span-based metrics that fail to capture what an adversary could actually infer, and assumes a single data subject, ignoring multi-subject scenarios. To address these limitations, we present…

Computation and Language · Computer Science 2026-04-24 Myeong Seok Oh , Dong-Yun Kim , Hanseok Oh , Chaean Kang , Joeun Kang , Xiaonan Wang , Hyunjung Park , Young Cheol Jung , Hansaem Kim

Most tasks in NLP require labeled data. Data labeling is often done on crowdsourcing platforms due to scalability reasons. However, publishing data on public platforms can only be done if no privacy-relevant information is included. Textual…

Computation and Language · Computer Science 2023-03-07 Nina Mouhammad , Johannes Daxenberger , Benjamin Schiller , Ivan Habernal

Fine-tuning Large Language Models (LLMs) on sensitive datasets carries a substantial risk of unintended memorization and leakage of Personally Identifiable Information (PII), which can violate privacy regulations and compromise individual…

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in processing and reasoning over diverse modalities, but their advanced abilities also raise significant privacy concerns, particularly regarding Personally…

Cryptography and Security · Computer Science 2025-10-01 Boyang Zhang , Istemi Ekin Akkus , Ruichuan Chen , Alice Dethise , Klaus Satzke , Ivica Rimac , Yang Zhang

The use of Natural Language Processing (NLP) in highstakes AI-based applications has increased significantly in recent years, especially since the emergence of Large Language Models (LLMs). However, despite their strong performance, LLMs…

We present PIIBench, a unified benchmark corpus for Personally Identifiable Information (PII) detection in natural language text. Existing resources for PII detection are fragmented across domain-specific corpora with mutually incompatible…

Computation and Language · Computer Science 2026-04-20 Pritesh Jha

Defining privacy and related notions such as Personal Identifiable Information (PII) is a central notion in computer science and other fields. The theoretical, technological, and application aspects of PII require a framework that provides…

Computers and Society · Computer Science 2018-03-28 Sabah S. Al-Fedaghi

The growing use of voice user interfaces has led to a surge in the collection and storage of speech data. While data collection allows for the development of efficient tools powering most speech services, it also poses serious privacy…

Cryptography and Security · Computer Science 2024-03-04 Pierre Champion

Text anonymization is the process of removing or obfuscating information from textual data to protect the privacy of individuals. This process inherently involves a complex trade-off between privacy protection and information preservation,…

Computation and Language · Computer Science 2025-09-23 Gabriel Loiseau , Damien Sileo , Damien Riquet , Maxime Meyer , Marc Tommasi

The increasing use of machine learning (ML) for Just-In-Time (JIT) defect prediction raises concerns about privacy leakage from software analytics data. Existing anonymization methods, such as tabular transformations and graph…

Software Engineering · Computer Science 2025-12-16 Maaz Khan , Gul Sher Khan , Ahsan Raza , Pir Sami Ullah , Abdul Ali Bangash

Artificial Intelligence (AI) faces growing challenges from evolving data protection laws and enforcement practices worldwide. Regulations like GDPR and CCPA impose strict compliance requirements on Machine Learning (ML) models, especially…

Machine Learning · Computer Science 2025-01-23 Shubhi Asthana , Ruchi Mahindru , Bing Zhang , Jorge Sanz

The growing reliance on large-scale speech data has made privacy protection a critical concern. However, existing anonymization approaches often degrade data utility, for example by disrupting acoustic continuity or reducing vocal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-21 Yunchong Xiao , Yuxiang Zhao , Ziyang Ma , Shuai Wang , Kai Yu , Jiachun Liao , Xie Chen

The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation…

Large language models (LLMs) exhibit powerful capabilities but risk memorizing sensitive personally identifiable information (PII) from their training data, posing significant privacy concerns. While machine unlearning techniques aim to…

Cryptography and Security · Computer Science 2026-01-23 Xinjie Zhou , Zhihui Yang , Lechao Cheng , Sai Wu , Gang Chen