中文
相关论文

相关论文: A Comparative Study of Light-weight Language Model…

200 篇论文

Privacy Masking is a critical concept under data privacy involving anonymization and de-anonymization of personally identifiable information (PII). Privacy masking techniques rely on Named Entity Recognition (NER) approaches under NLP…

计算与语言 · 计算机科学 2025-04-18 Devansh Singh , Sundaraparipurnan Narayanan

The widespread adoption of Large Language Models (LLMs) has raised significant privacy concerns regarding the exposure of personally identifiable information (PII) in user prompts. To address this challenge, we propose a query-unrelated PII…

密码学与安全 · 计算机科学 2026-02-18 Hao Shen , Zhouhong Gu , Haokai Hong , Weili Han

Detecting personally identifiable information (PII) in user queries is critical for ensuring privacy in question-answering systems. Current approaches mainly redact all PII, disregarding the fact that some of them may be contextually…

密码学与安全 · 计算机科学 2026-02-11 Mariia Ponomarenko , Sepideh Abedini , Masoumeh Shafieinejad , D. B. Emerson , Shubhankar Mohapatra , Xi He

Removing Personally Identifiable Information (PII) from clinical notes in Electronic Health Records (EHRs) is essential for research and AI development. While Large Language Models (LLMs) are powerful, their high computational costs and the…

计算与语言 · 计算机科学 2025-10-23 Prakrithi Shivaprakash , Lekhansh Shukla , Animesh Mukherjee , Prabhat Chand , Pratima Murthy

The detection of Personally Identifiable Information (PII) is critical for privacy compliance but remains challenging in low-resource languages due to linguistic diversity and limited annotated data. We present RECAP, a hybrid framework…

Fine-tuning Large Language Models (LLMs) on sensitive datasets carries a substantial risk of unintended memorization and leakage of Personally Identifiable Information (PII), which can violate privacy regulations and compromise individual…

Language Models (LMs) have been shown to leak information about training data through sentence-level membership inference and reconstruction attacks. Understanding the risk of LMs leaking Personally Identifiable Information (PII) has…

机器学习 · 计算机科学 2023-04-25 Nils Lukas , Ahmed Salem , Robert Sim , Shruti Tople , Lukas Wutschitz , Santiago Zanella-Béguelin

Redacting Personally Identifiable Information (PII) from unstructured text is critical for ensuring data privacy in regulated domains. While earlier approaches have relied on rule-based systems and domain-specific Named Entity Recognition…

密码学与安全 · 计算机科学 2025-08-08 Leon Garza , Anantaa Kotal , Aritran Piplai , Lavanya Elluri , Prajit Das , Aman Chadha

Reliable detection of personally identifiable information (PII) is increasingly important across modern data-processing systems, yet the task remains difficult: PII spans are heterogeneous, locale-dependent, context-sensitive, and often…

计算与语言 · 计算机科学 2026-05-12 Urchade Zaratiana , Ash Lewis , George Hurn-Maloney

AI chatbots have quietly become the world's most popular therapists, coaches, and confidants. Users of cloud-based LLM services are increasingly shifting from simple queries like idea generation and poem writing, to deeply personal…

人机交互 · 计算机科学 2026-04-08 Max Holschneider , Saetbyeol LeeYouk

The advancement of large language models (LLMs) brings notable improvements across various applications, while simultaneously raising concerns about potential private data exposure. One notable capability of LLMs is their ability to form…

计算与语言 · 计算机科学 2024-02-12 Hanyin Shao , Jie Huang , Shen Zheng , Kevin Chen-Chuan Chang

Artificial Intelligence (AI) faces growing challenges from evolving data protection laws and enforcement practices worldwide. Regulations like GDPR and CCPA impose strict compliance requirements on Machine Learning (ML) models, especially…

机器学习 · 计算机科学 2025-01-23 Shubhi Asthana , Ruchi Mahindru , Bing Zhang , Jorge Sanz

Protecting sensitive information is crucial in today's world of Large Language Models (LLMs) and data-driven services. One common method used to preserve privacy is by using data perturbation techniques to reduce overreaching utility of…

计算与语言 · 计算机科学 2023-07-19 Ajinkya Deshmukh , Saumya Banthia , Anantha Sharma

Personally identifiable information (PII) anonymization is a high-stakes task that poses a barrier to many open-science data sharing initiatives. While PII identification has made large strides in recent years, in practice, error thresholds…

计算与语言 · 计算机科学 2025-05-23 Matthew Zent , Digory Smith , Simon Woodhead

Many datasets contain personally identifiable information, or PII, which poses privacy risks to individuals. PII masking is commonly used to redact personal information such as names, addresses, and phone numbers from text data. Most modern…

计算与语言 · 计算机科学 2022-05-11 Courtney Mansfield , Amandalynne Paullada , Kristen Howell

Speaker, author, and other biometric identification applications often compare a sample's similarity to a database of templates to determine the identity. Given that data may be noisy and similarity measures can be inaccurate, such a…

音频与语音处理 · 电气工程与系统科学 2025-12-01 Tom Bäckström , Mohammad Hassan Vali , My Nguyen , Silas Rech

As Large Language Models (LLMs) gain wider adoption, ensuring their reliable handling of Personally Identifiable Information (PII) across diverse regulatory contexts has become essential. This work introduces a scalable multilingual data…

The current literature on memorization in Natural Language Models, especially Large Language Models (LLMs), poses severe security and privacy risks, as models tend to memorize personally identifying information (PIIs) from training data. We…

计算与语言 · 计算机科学 2026-02-19 Kunj Joshi , David A. Smith

Large language models (LLMs) are renowned for their extensive linguistic knowledge and strong generalization capabilities, but their high computational demands make them unsuitable for resource-constrained environments. In contrast, small…

计算与语言 · 计算机科学 2025-06-10 Kyeonghyun Kim , Jinhee Jang , Juhwan Choi , Yoonji Lee , Kyohoon Jin , YoungBin Kim

Concerns regarding Large Language Models (LLMs) to memorize and disclose private information, particularly Personally Identifiable Information (PII), become prominent within the community. Many efforts have been made to mitigate the privacy…

机器学习 · 计算机科学 2024-05-21 Ruizhe Chen , Tianxiang Hu , Yang Feng , Zuozhu Liu
‹ 上一页 1 2 3 10 下一页 ›