English
Related papers

Related papers: Subject-level Inference for Realistic Text Anonymi…

200 papers

The current privacy evaluation for speaker anonymization often overestimates privacy when a same-gender target selection algorithm (TSA) is used, although this TSA leaks the speaker's gender and should hence be more vulnerable. We…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Carlos Franzreb , Arnab Das , Tim Polzehl , Sebastian Möller

De-identification of data used for automatic speech recognition modeling is a critical component in protecting privacy, especially in the medical domain. However, simply removing all personally identifiable information (PII) from end-to-end…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-13 Martin Flechl , Shou-Chun Yin , Junho Park , Peter Skala

LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private. Existing systems support only suppression (omitting sensitive information) and…

Cryptography and Security · Computer Science 2026-04-09 Yunze Xiao , Wenkai Li , Xiaoyuan Wu , Ningshan Ma , Yueqi Song , Weihao Xuan

The proliferation of textual data containing sensitive personal information across various domains requires robust anonymization techniques to protect privacy and comply with regulations, while preserving data usability for diverse and…

Computation and Language · Computer Science 2025-12-17 Tobias Deußer , Lorenz Sparrenberg , Armin Berger , Max Hahnbück , Christian Bauckhage , Rafet Sifa

The growing use of voice user interfaces has led to a surge in the collection and storage of speech data. While data collection allows for the development of efficient tools powering most speech services, it also poses serious privacy…

Cryptography and Security · Computer Science 2024-03-04 Pierre Champion

For sensitive text data to be shared among NLP researchers and practitioners, shared documents need to comply with data protection and privacy laws. There is hence a growing interest in automated approaches for text anonymization. However,…

Computation and Language · Computer Science 2021-03-18 Maximilian Mozes , Bennett Kleinberg

The growing reliance on large-scale speech data has made privacy protection a critical concern. However, existing anonymization approaches often degrade data utility, for example by disrupting acoustic continuity or reducing vocal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-21 Yunchong Xiao , Yuxiang Zhao , Ziyang Ma , Shuai Wang , Kai Yu , Jiachun Liao , Xie Chen

Adolescent suicide is a critical global health issue, and speech provides a cost-effective modality for automatic suicide risk detection. Given the vulnerable population, protecting speaker identity is particularly important, as speech…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Ziyun Cui , Sike Jia , Yang Lin , Yinan Duan , Diyang Qu , Runsen Chen , Chao Zhang , Chang Lei , Wen Wu

In recent years, there has been an increasing interest in image anonymization, particularly focusing on the de-identification of faces and individuals. However, for self-driving applications, merely de-identifying faces and individuals…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Dongyu Liu , Xuhong Wang , Cen Chen , Yanhao Wang , Shengyue Yao , Yilun Lin

The purpose of anonymizing structured data is to protect the privacy of individuals in the data while retaining the statistical properties of the data. An important class of attack on anonymized data is attribute inference, where an…

Cryptography and Security · Computer Science 2025-07-03 Paul Francis , David Wagner

Large Language Models (LLMs) are trained on large-scale web data, which makes it difficult to grasp the contribution of each text. This poses the risk of leaking inappropriate data such as benchmarks, personal information, and copyrighted…

Computation and Language · Computer Science 2024-04-18 Masahiro Kaneko , Youmi Ma , Yuki Wata , Naoaki Okazaki

Automated masking of Personally Identifiable Information (PII) is critical for privacy-preserving conversational systems. While current frontier large language models demonstrate strong PII masking capabilities, concerns about data handling…

Computation and Language · Computer Science 2025-12-23 Prabigya Acharya , Liza Shrestha

Accurate recognition of personally identifiable information (PII) is central to automated text anonymization. This paper investigates the effectiveness of cross-domain model transfer, multi-domain data fusion, and sample-efficient learning…

Computation and Language · Computer Science 2026-01-13 Junhong Ye , Xu Yuan , Xinying Qiu

Removing personally identifiable information (PII) from texts is necessary to comply with various data protection regulations and to enable data sharing without compromising privacy. However, recent works show that documents sanitized by…

Computation and Language · Computer Science 2026-03-16 Sebastian Ochs , Ivan Habernal

Private Set Intersection (PSI) is a widely used protocol that enables two parties to securely compute a function over the intersected part of their shared datasets and has been a significant research focus over the years. However, recent…

Cryptography and Security · Computer Science 2023-12-01 Bo Jiang , Jian Du , Qiang Yan

Data sanitization in the context of language modeling involves identifying sensitive content, such as personally identifiable information (PII), and redacting them from a dataset corpus. It is a common practice used in natural language…

Computation and Language · Computer Science 2024-11-12 Anwesan Pal , Radhika Bhargava , Kyle Hinsz , Jacques Esterhuizen , Sudipta Bhattacharya

Named Entity Recognition is the task to locate and classify the entities in the text. However, Unlabeled Entity Problem in NER datasets seriously hinders the improvement of NER performance. This paper proposes SCL-RAI to cope with this…

Computation and Language · Computer Science 2023-10-25 Shuzheng Si , Shuang Zeng , Jiaxing Lin , Baobao Chang

Deidentification seeks to anonymize textual data prior to distribution. Automatic deidentification primarily uses supervised named entity recognition from human-labeled data points. We propose an unsupervised deidentification method that…

Computation and Language · Computer Science 2022-10-24 John X. Morris , Justin T. Chiu , Ramin Zabih , Alexander M. Rush

We propose a novel method to bootstrap text anonymization models based on distant supervision. Instead of requiring manually labeled training data, the approach relies on a knowledge graph expressing the background information assumed to be…

Computation and Language · Computer Science 2022-05-17 Anthi Papadopoulou , Pierre Lison , Lilja Øvrelid , Ildikó Pilán

Anonymizing textual documents is a highly context-sensitive problem: the appropriate balance between privacy protection and utility preservation varies with the data domain, privacy objectives, and downstream application. However, existing…

Computation and Language · Computer Science 2026-04-21 Gabriel Loiseau , Damien Sileo , Damien Riquet , Maxime Meyer , Marc Tommasi