English
Related papers

Related papers: DIRI: Adversarial Patient Reidentification with La…

200 papers

Data containing personal information is increasingly used to train, fine-tune, or query Large Language Models (LLMs). Text is typically scrubbed of identifying information prior to use, often with tools such as Microsoft's Presidio or…

Computation and Language · Computer Science 2026-02-16 Nataša Krčo , Zexi Yao , Matthieu Meeus , Yves-Alexandre de Montjoye

De-identification in the healthcare setting is an application of NLP where automated algorithms are used to remove personally identifying information of patients (and, sometimes, providers). With the recent rise of generative large language…

Computation and Language · Computer Science 2025-09-19 Kiana Aghakasiri , Noopur Zambare , JoAnn Thai , Carrie Ye , Mayur Mehta , J. Ross Mitchell , Mohamed Abdalla

Large language models (LLMs) have shown strong performance on clinical de-identification, the task of identifying sensitive identifiers to protect privacy. However, previous work has not examined their generalizability between formats,…

Computation and Language · Computer Science 2026-02-19 Noopur Zambare , Kiana Aghakasiri , Carissa Lin , Carrie Ye , J. Ross Mitchell , Mohamed Abdalla

De-identification of clinical text remains essential for secondary use of electronic health records (EHRs), yet public benchmarks such as i2b2 2006/2014 are over a decade old and lack the semantic and demographic diversity of modern…

Computation and Language · Computer Science 2026-05-06 Jose D. Posada , David Love , Somalee Datta , Priya Desai

Background : De-identification of DICOM (Digital Imaging and Communi-cations in Medicine) files is an essential component of medical image research. Personal Identifiable Information (PII) and/or Personal Health Identifying Information…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Bufano Michele , Kotter Elmar

Removing Personally Identifiable Information (PII) from clinical notes in Electronic Health Records (EHRs) is essential for research and AI development. While Large Language Models (LLMs) are powerful, their high computational costs and the…

Computation and Language · Computer Science 2025-10-23 Prakrithi Shivaprakash , Lekhansh Shukla , Animesh Mukherjee , Prabhat Chand , Pratima Murthy

Sharing clinical research data is key for increasing the pace of medical discoveries that improve human health. However, concern about study participants' privacy, confidentiality, and safety is a major factor that deters researchers from…

Automated clinical text anonymization has the potential to unlock the widespread sharing of textual health data for secondary usage while assuring patient privacy and safety. Despite the proposal of many complex and theoretically successful…

Computation and Language · Computer Science 2024-06-05 David Pissarra , Isabel Curioso , João Alveira , Duarte Pereira , Bruno Ribeiro , Tomás Souper , Vasco Gomes , André V. Carreiro , Vitor Rolla

Unstructured textual data are at the heart of health systems: liaison letters between doctors, operating reports, coding of procedures according to the ICD-10 standard, etc. The details included in these documents make it possible to get to…

Cryptography and Security · Computer Science 2023-10-09 Yakini Tchouka , Jean-François Couchot , Maxime Coulmeau , David Laiymani , Philippe Selles , Azzedine Rahmani

Ensuring clinical data privacy while preserving utility is critical for AI-driven healthcare and data analytics. Existing de-identification (De-ID) methods, including rule-based techniques, deep learning models, and large language models…

Artificial Intelligence · Computer Science 2025-07-28 Praphul Singh , Charlotte Dzialo , Jangwon Kim , Sumana Srivatsa , Irfan Bulu , Sri Gadde , Krishnaram Kenthapadi

Confidentiality of patient information is an essential part of Electronic Health Record System. Patient information, if exposed, can cause a serious damage to the privacy of individuals receiving healthcare. Hence it is important to remove…

Machine Learning · Computer Science 2018-08-06 Abhai Kollara Dilip , Kamal Raj K , Malaikannan Sankarasubbu

The digitization of healthcare has facilitated the sharing and re-using of medical data but has also raised concerns about confidentiality and privacy. HIPAA (Health Insurance Portability and Accountability Act) mandates removing…

The objective of this study is to address the critical issue of de-identification of clinical reports in order to allow access to data for research purposes, while ensuring patient privacy. The study highlights the difficulties faced in…

Computation and Language · Computer Science 2023-03-24 Xavier Tannier , Perceval Wajsbürt , Alice Calliger , Basile Dura , Alexandre Mouchet , Martin Hilka , Romain Bey

Unstructured information in electronic health records provide an invaluable resource for medical research. To protect the confidentiality of patients and to conform to privacy regulations, de-identification methods automatically remove…

Computation and Language · Computer Science 2020-01-17 Jan Trienes , Dolf Trieschnigg , Christin Seifert , Djoerd Hiemstra

Background: More than half (57%) of pharma clinical research spend is in support of clinical trials. One reason is that Electronic Health Record (EHR) systems and HIPAA privacy rules often limit how broadly patient information can be…

Quantitative Methods · Quantitative Biology 2018-07-03 Andrew J McMurry , Richen Zhang , Alex Foxman , Lawrence Reiter , Ronny Schnel , DeLeys Brandman

Electronic Health Records (EHRs) have become the primary form of medical data-keeping across the United States. Federal law restricts the sharing of any EHR data that contains protected health information (PHI). De-identification, the…

Computation and Language · Computer Science 2021-03-26 Abdullah Ahmed , Adeel Abbasi , Carsten Eickhoff

Objectives; The accumulation and usefulness of clinical data have increased with IT development. While using clinical data that needs to be identifiable to obtain meaningful information, it is essential to ensure that data is de-identified…

Cryptography and Security · Computer Science 2018-04-16 Jipmin Jung , Phillip Park , Jaedong Lee , Hyein Lee , Geonkook Lee , Hyosoung Cha

De-identification is the task of identifying protected health information (PHI) in the clinical text. Existing neural de-identification models often fail to generalize to a new dataset. We propose a simple yet effective data augmentation…

Computation and Language · Computer Science 2020-10-13 Xiang Yue , Shuang Zhou

Massive digital data processing provides a wide range of opportunities and benefits, but at the cost of endangering personal data privacy. Anonymisation consists in removing or replacing sensitive information from data, enabling its…

Computation and Language · Computer Science 2020-03-18 Aitor García-Pablos , Naiara Perez , Montse Cuadros

Large language models (LLMs) excel at clinical information extraction but their computational demands limit practical deployment. Knowledge distillation--the process of transferring knowledge from larger to smaller models--offers a…

Computation and Language · Computer Science 2025-01-03 Karthik S. Vedula , Annika Gupta , Akshay Swaminathan , Ivan Lopez , Suhana Bedi , Nigam H. Shah