English
Related papers

Related papers: GLiNER2-PII: A Multilingual Model for Personally I…

200 papers

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing but also pose significant privacy risks by memorizing and leaking Personally Identifiable Information (PII). Existing mitigation…

Machine Learning · Computer Science 2025-03-17 Ahmed Frikha , Muhammad Reza Ar Razi , Krishna Kanth Nakka , Ricardo Mendes , Xue Jiang , Xuebing Zhou

The memorization of sensitive and personally identifiable information (PII) by large language models (LLMs) poses growing privacy risks as models scale and are increasingly deployed in real-world applications. Existing efforts to study…

Computation and Language · Computer Science 2025-05-20 Sriram Selvam , Anneswa Ghosh

The rapid advancement and widespread use of large language models (LLMs) have raised significant concerns regarding the potential leakage of personally identifiable information (PII). These models are often trained on vast quantities of…

Cryptography and Security · Computer Science 2023-07-06 Siwon Kim , Sangdoo Yun , Hwaran Lee , Martin Gubri , Sungroh Yoon , Seong Joon Oh

The advancement of large language models (LLMs) brings notable improvements across various applications, while simultaneously raising concerns about potential private data exposure. One notable capability of LLMs is their ability to form…

Computation and Language · Computer Science 2024-02-12 Hanyin Shao , Jie Huang , Shen Zheng , Kevin Chen-Chuan Chang

Redacting Personally Identifiable Information (PII) from unstructured text is critical for ensuring data privacy in regulated domains. While earlier approaches have relied on rule-based systems and domain-specific Named Entity Recognition…

Cryptography and Security · Computer Science 2025-08-08 Leon Garza , Anantaa Kotal , Aritran Piplai , Lavanya Elluri , Prajit Das , Aman Chadha

Many datasets contain personally identifiable information, or PII, which poses privacy risks to individuals. PII masking is commonly used to redact personal information such as names, addresses, and phone numbers from text data. Most modern…

Computation and Language · Computer Science 2022-05-11 Courtney Mansfield , Amandalynne Paullada , Kristen Howell

Users interacting with large language models (LLMs) under their real identifiers often unknowingly risk disclosing private information. Automatically notifying users whether their queries leak privacy and which phrases leak what private…

Computation and Language · Computer Science 2025-08-11 Hang Zeng , Xiangyu Liu , Yong Hu , Chaoyue Niu , Fan Wu , Shaojie Tang , Guihai Chen

Large Language Models (LLMs) excel in various domains but pose inherent privacy risks. Existing methods to evaluate privacy leakage in LLMs often use memorized prefixes or simple instructions to extract data, both of which well-alignment…

Cryptography and Security · Computer Science 2025-05-19 Yidan Wang , Yanan Cao , Yubing Ren , Fang Fang , Zheng Lin , Binxing Fang

The rise of chronic diseases and pandemics like COVID-19 has emphasized the need for effective patient data processing while ensuring privacy through anonymization and de-identification of protected health information (PHI). Anonymized data…

Computation and Language · Computer Science 2024-12-17 Murat Gunay , Bunyamin Keles , Raife Hizlan

Large Language Models (LLMs) hold promise for advancing legal practice by automating complex tasks and improving access to justice. However, their adoption is limited by concerns over client confidentiality, especially when lawyers include…

Computation and Language · Computer Science 2025-01-22 M. Mikail Demir , Hakan T. Otal , M. Abdullah Canbaz

Defining privacy and related notions such as Personal Identifiable Information (PII) is a central notion in computer science and other fields. The theoretical, technological, and application aspects of PII require a framework that provides…

Computers and Society · Computer Science 2018-03-28 Sabah S. Al-Fedaghi

The widespread usage of large-scale multimodal models like CLIP has heightened concerns about the leakage of PII. Existing methods for identity inference in CLIP models require querying the model with full PII, including textual…

Machine Learning · Computer Science 2025-03-26 Songze Li , Ruoxi Cheng , Xiaojun Jia

Machine learning practitioners often fine-tune generative pre-trained models like GPT-3 to improve model performance at specific tasks. Previous works, however, suggest that fine-tuned machine learning models memorize and emit sensitive…

Machine Learning · Computer Science 2024-04-17 Albert Yu Sun , Eliott Zemour , Arushi Saxena , Udith Vaidyanathan , Eric Lin , Christian Lau , Vaikkunth Mugunthan

Crash narratives in crash reports provide crucial contextual information for traffic safety analysis. Yet, their broader use is hindered by the presence of personally identifiable information (PII), including names, home addresses, and…

Cryptography and Security · Computer Science 2026-04-20 Junyi Ma , Pei Li , Rui Gan , Kai Cheng , Steven T. Parker , Bin Ran

With the increasing use of conversational AI systems, there is growing concern over privacy leaks, especially when users share sensitive personal data in interactions with Large Language Models (LLMs). Conversations shared with these models…

Computation and Language · Computer Science 2025-11-03 Jayden Serenari , Stephen Lee

De-identification of data used for automatic speech recognition modeling is a critical component in protecting privacy, especially in the medical domain. However, simply removing all personally identifiable information (PII) from end-to-end…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-13 Martin Flechl , Shou-Chun Yin , Junho Park , Peter Skala

The ever-increasing adoption of Large Language Models in critical sectors like finance, healthcare, and government raises privacy concerns regarding the handling of sensitive Personally Identifiable Information (PII) during training. In…

Machine Learning · Computer Science 2026-01-06 Intae Jeon , Yujeong Kwon , Hyungjoon Koo

Efficiently selecting relevant content from vast candidate pools is a critical challenge in modern recommender systems. Traditional methods, such as item-to-item collaborative filtering (CF) and two-tower models, often fall short in…

Information Retrieval · Computer Science 2026-01-26 Shaoqing Wang , Yingcai Ma , Kairui Fu , Ziyang Wang , Dunxian Huang , Yuliang Yan , Jian Wu

Collecting personally identifiable information (PII) on data subjects has become big business. Data brokers and data processors are part of a multi-billion-dollar industry that profits from collecting, buying, and selling consumer data. Yet…

Computation and Language · Computer Science 2023-02-20 Juniper Lovato , Philip Mueller , Parisa Suchdev , Peter S. Dodds

Large language models (LLMs) have transformed natural language processing, but their ability to memorize training data poses significant privacy risks. This paper investigates model inversion attacks on the Llama 3.2 model, a multilingual…

Machine Learning · Computer Science 2025-07-08 Sathesh P. Sivashanmugam