English
Related papers

Related papers: GLiNER2-PII: A Multilingual Model for Personally I…

200 papers

Large Language Models (LLMs) pose significant privacy risks, potentially leaking training data due to implicit memorization. Existing privacy attacks primarily focus on membership inference attacks (MIAs) or data extraction attacks, but…

Computation and Language · Computer Science 2025-06-11 Wenlong Meng , Zhenyuan Guo , Lenan Wu , Chen Gong , Wenyan Liu , Weixian Li , Chengkun Wei , Wenzhi Chen

Much text describes a changing world (e.g., procedures, stories, newswires), and understanding them requires tracking how entities change. An earlier dataset, OpenPI, provided crowdsourced annotations of entity state changes in text.…

Computation and Language · Computer Science 2024-01-26 Li Zhang , Hainiu Xu , Abhinav Kommula , Chris Callison-Burch , Niket Tandon

Accurate recognition of personally identifiable information (PII) is central to automated text anonymization. This paper investigates the effectiveness of cross-domain model transfer, multi-domain data fusion, and sample-efficient learning…

Computation and Language · Computer Science 2026-01-13 Junhong Ye , Xu Yuan , Xinying Qiu

Due to the sensitive nature of personally identifiable information (PII), its owners may have the authority to control its inclusion or request its removal from large-language model (LLM) training. Beyond this, PII may be added or removed…

Computation and Language · Computer Science 2025-06-27 Jaydeep Borkar , Matthew Jagielski , Katherine Lee , Niloofar Mireshghallah , David A. Smith , Christopher A. Choquette-Choo

Text de-identification techniques are often used to mask personally identifiable information (PII) from documents. Their ability to conceal the identity of the individuals mentioned in a text is, however, hard to measure. Recent work has…

Computation and Language · Computer Science 2025-10-13 Lucas Georges Gabriel Charpentier , Pierre Lison

In this work, we introduce PII-Scope, a comprehensive benchmark designed to evaluate state-of-the-art methodologies for PII extraction attacks targeting LLMs across diverse threat settings. Our study provides a deeper understanding of these…

Computation and Language · Computer Science 2025-05-27 Krishna Kanth Nakka , Ahmed Frikha , Ricardo Mendes , Xue Jiang , Xuebing Zhou

Sanitizing sensitive text data typically involves removing personally identifiable information (PII) or generating synthetic data under the assumption that these methods adequately protect privacy; however, their effectiveness is often only…

Cryptography and Security · Computer Science 2026-03-17 Rui Xin , Niloofar Mireshghallah , Shuyue Stella Li , Michael Duan , Hyunwoo Kim , Yejin Choi , Yulia Tsvetkov , Sewoong Oh , Pang Wei Koh

Prompt serves as a crucial link in interacting with large language models (LLMs), widely impacting the accuracy and interpretability of model outputs. However, acquiring accurate and high-quality responses necessitates precise prompts,…

Cryptography and Security · Computer Science 2024-08-20 Xiongtao Sun , Gan Liu , Zhipeng He , Hui Li , Xiaoguang Li

Information extraction tasks require both accurate, efficient, and generalisable models. Classical supervised deep learning approaches can achieve the required performance, but they need large datasets and are limited in their ability to…

Machine Learning · Computer Science 2024-08-02 Ihor Stepanov , Mykhailo Shtopko

Objectives: Despite the recent adoption of large language models (LLMs) for biomedical information extraction, challenges in prompt engineering and algorithms persist, with no dedicated software available. To address this, we developed…

Machine Learning · Computer Science 2025-04-02 Enshuo Hsu , Kirk Roberts

Federated large language models (FedLLMs) enable cross-silo collaborative training among institutions while preserving data locality, making them appealing for privacy-sensitive domains such as law, finance, and healthcare. However, the…

Computation and Language · Computer Science 2026-02-26 Yingqi Hu , Zhuo Zhang , Jingyuan Zhang , Jinghua Wang , Qifan Wang , Lizhen Qu , Zenglin Xu

The deployment of Large Language Models in agentic, multi-turn conversational settings has introduced a class of privacy vulnerabilities that existing protection mechanisms are not designed to address. Current approaches to Personally…

Cryptography and Security · Computer Science 2026-04-21 Aman Panjwani

Production LLM systems require both safety moderation and PII detection under strict latency and cost constraints. This creates a trade-off: autoregressive moderators are accurate but expensive, while lightweight encoders are faster but…

Cryptography and Security · Computer Science 2026-05-08 Bogdan Minko , Sabrina Sadiekh , Evgeniy Kokuykin

With the rise of deep learning in various applications, privacy concerns around the protection of training data have become a critical area of research. Whereas prior studies have focused on privacy risks in single-modal models, we…

Machine Learning · Computer Science 2024-07-10 Dominik Hintersdorf , Lukas Struppek , Manuel Brack , Felix Friedrich , Patrick Schramowski , Kristian Kersting

Despite advances in the use of large language models (LLMs) in downstream tasks, their ability to memorize information has raised privacy concerns. Therefore, protecting personally identifiable information (PII) during LLM training remains…

Machine Learning · Computer Science 2025-12-02 Stella Etuk , Ashraf Matrawy

Measuring personal disclosures made in human-chatbot interactions can provide a better understanding of users' AI literacy and facilitate privacy research for large language models (LLMs). We run an extensive, fine-grained analysis on the…

Computation and Language · Computer Science 2024-07-23 Niloofar Mireshghallah , Maria Antoniak , Yash More , Yejin Choi , Golnoosh Farnadi

With the rise of large language models (LLMs), increasing research has recognized their risk of leaking personally identifiable information (PII) under malicious attacks. Although efforts have been made to protect PII in LLMs, existing…

Computation and Language · Computer Science 2025-03-12 Martin Kuo , Jingyang Zhang , Jianyi Zhang , Minxue Tang , Louis DiValentin , Aolin Ding , Jingwei Sun , William Chen , Amin Hass , Tianlong Chen , Yiran Chen , Hai Li

Biomedical named entity recognition (NER) presents unique challenges due to specialized vocabularies, the sheer volume of entities, and the continuous emergence of novel entities. Traditional NER models, constrained by fixed taxonomies and…

Computation and Language · Computer Science 2025-05-22 Anthony Yazdani , Ihor Stepanov , Douglas Teodoro

Self-disclosure, while being common and rewarding in social media interaction, also poses privacy risks. In this paper, we take the initiative to protect the user-side privacy associated with online self-disclosure through detection and…

Computation and Language · Computer Science 2024-06-25 Yao Dou , Isadora Krsek , Tarek Naous , Anubha Kabra , Sauvik Das , Alan Ritter , Wei Xu

Temporal expression identification is crucial for understanding texts written in natural language. Although highly effective systems such as HeidelTime exist, their limited runtime performance hampers adoption in large-scale applications…

Computation and Language · Computer Science 2024-03-26 Hugo Sousa , Ricardo Campos , Alípio Jorge