English
Related papers

Related papers: PIIvot: A Lightweight NLP Anonymization Framework …

200 papers

We present MimiTalk, a dual-agent constitutional AI framework designed for scalable and ethical conversational data collection in social science research. The framework integrates a supervisor model for strategic oversight and a…

Human-Computer Interaction · Computer Science 2025-11-07 Fengming Liu , Shubin Yu

Despite advances in the use of large language models (LLMs) in downstream tasks, their ability to memorize information has raised privacy concerns. Therefore, protecting personally identifiable information (PII) during LLM training remains…

Machine Learning · Computer Science 2025-12-02 Stella Etuk , Ashraf Matrawy

Speaker anonymization aims to conceal a speaker's identity without degrading speech quality and intelligibility. Most speaker anonymization systems disentangle the speaker representation from the original speech and achieve anonymization by…

Sound · Computer Science 2023-10-10 Yuanjun Lv , Jixun Yao , Peikun Chen , Hongbin Zhou , Heng Lu , Lei Xie

Clinical text processing has gained more and more attention in recent years. The access to sensitive patient data, on the other hand, is still a big challenge, as text cannot be shared without legal hurdles and without removing personal…

Computation and Language · Computer Science 2022-09-02 Iyadh Ben Cheikh Larbi , Aljoscha Burchardt , Roland Roller

Large Language Models (LLMs) memorize, and thus, among huge amounts of uncontrolled data, may memorize Personally Identifiable Information (PII), which should not be stored and, consequently, not leaked. In this paper, we introduce Private…

Cryptography and Security · Computer Science 2025-08-22 Elena Sofia Ruzzetti , Giancarlo A. Xompero , Davide Venditti , Fabio Massimo Zanzotto

Large language model (LLM) agents are increasingly deployed in personalized tasks involving sensitive, context-dependent information, where privacy violations may arise in agents' action due to the implicitness of contextual privacy.…

Computation and Language · Computer Science 2026-02-17 Yuhan Cheng , Hancheng Ye , Hai Helen Li , Jingwei Sun , Yiran Chen

Privacy-preserving voice protection approaches primarily suppress privacy-related information derived from paralinguistic attributes while preserving the linguistic content. Existing solutions focus particularly on single-speaker scenarios.…

Sound · Computer Science 2025-03-28 Xiaoxiao Miao , Ruijie Tao , Chang Zeng , Xin Wang

Small language models (SLMs) become unprecedentedly appealing due to their approximately equivalent performance compared to large language models (LLMs) in certain fields with less energy and time consumption during training and inference.…

Computation and Language · Computer Science 2025-09-29 Jieli Zhu , Vi Ngoc-Nha Tran

Federated learning (FL) enables collaborative training across organizational silos without sharing raw data, making it attractive for privacy-sensitive applications. With the rapid adoption of large language models (LLMs), federated…

Cryptography and Security · Computer Science 2026-04-17 Mohamed Shaaban , Mohamed Elmahallawy

This paper introduces a conversational interface system that enables participatory design of differentially private AI systems in public sector applications. Addressing the challenge of balancing mathematical privacy guarantees with…

Information Theory · Computer Science 2025-05-28 Wenjun Yang , Eyhab Al-Masri

Face anonymization aims to conceal the visual identity of a face to safeguard the individual's privacy. Traditional methods like blurring and pixelation can largely remove identifying features, but these techniques significantly degrade…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Lin Yuan , Kai Liang , Xiong Li , Tao Wu , Nannan Wang , Xinbo Gao

Pre-trained Large Language Models (LLMs) are an integral part of modern AI that have led to breakthrough performances in complex AI tasks. Major AI companies with expensive infrastructures are able to develop and train these large models…

Cryptography and Security · Computer Science 2023-05-02 Rouzbeh Behnia , Mohamamdreza Ebrahimi , Jason Pacheco , Balaji Padmanabhan

Large Language Models (LLMs) are gaining increasing attention due to their exceptional performance across numerous tasks. As a result, the general public utilize them as an influential tool for boosting their productivity while natural…

Cryptography and Security · Computer Science 2023-06-16 Zhigang Kan , Linbo Qiao , Hao Yu , Liwen Peng , Yifu Gao , Dongsheng Li

Text sanitization is the task of redacting a document to mask all occurrences of (direct or indirect) personal identifiers, with the goal of concealing the identity of the individual(s) referred in it. In this paper, we consider a two-step…

Computation and Language · Computer Science 2023-10-24 Anthi Papadopoulou , Pierre Lison , Mark Anderson , Lilja Øvrelid , Ildikó Pilán

Large Language Models (LLMs) have been reported to "leak" Personally Identifiable Information (PII), with successful PII reconstruction often interpreted as evidence of memorization. We propose a principled revision of memorization…

Computation and Language · Computer Science 2026-01-08 Xiaoyu Luo , Yiyi Chen , Qiongxiu Li , Johannes Bjerva

The proliferation of Large Language Models (LLMs) has driven considerable interest in fine-tuning them with domain-specific data to create specialized language models. Nevertheless, such domain-specific fine-tuning data often contains…

Computation and Language · Computer Science 2024-10-29 Yijia Xiao , Yiqiao Jin , Yushi Bai , Yue Wu , Xianjun Yang , Xiao Luo , Wenchao Yu , Xujiang Zhao , Yanchi Liu , Quanquan Gu , Haifeng Chen , Wei Wang , Wei Cheng

Knowledge graph-based dialogue generation (KG-DG) is a challenging task requiring models to effectively incorporate external knowledge into conversational responses. While large language models (LLMs) have achieved impressive results across…

Computation and Language · Computer Science 2025-11-18 Hadi Sheikhi , Chenyang Huang , Osmar R. Zaïane

The trend of scaling up speech generation models poses a threat of biometric information leakage of the identities of the voices in the training data, raising privacy and security concerns. In this paper, we investigate training…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-21 Wen-Chin Huang , Yi-Chiao Wu , Tomoki Toda

Artificial intelligence systems increasingly mediate knowledge, communication, and decision making. Development and governance remain concentrated within a small set of firms and states, raising concerns that technologies may encode narrow…

Artificial Intelligence · Computer Science 2026-02-12 Rashid Mushkani

The widespread availability of large-scale code datasets has fueled the rapid development of large language models (LLMs) for code-related tasks. These datasets may include sensitive personally identifiable information (PII), which can lead…

Software Engineering · Computer Science 2026-05-18 Yifei Ge , Zhenpeng Chen , Weisong Sun , Yuchen Chen , Chunrong Fang , Juan Zhai , Xiaofang Zhang , Xia Feng , Yang Liu , Zhenyu Chen