English
Related papers

Related papers: DeIDClinic: A Risk-Aware Pseudonymization Framewor…

200 papers

Clinical free-text data offers immense potential to improve population health research such as richer phenotyping, symptom tracking, and contextual understanding of patient care. However, these data present significant privacy risks due to…

Free-text clinical notes detail all aspects of patient care and have great potential to facilitate quality improvement and assurance initiatives as well as advance clinical research. However, concerns about patient privacy and…

Computation and Language · Computer Science 2021-02-23 Nicholas Dobbins , David Wayne , Kahyun Lee , Özlem Uzuner , Meliha Yetisgen

Unstructured textual data is at the heart of healthcare systems. For obvious privacy reasons, these documents are not accessible to researchers as long as they contain personally identifiable information. One way to share this data while…

Cryptography and Security · Computer Science 2022-11-03 Yakini Tchouka , Jean-François Couchot , David Laiymani

Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporary deidentification…

Computation and Language · Computer Science 2024-10-23 John X. Morris , Thomas R. Campion , Sri Laasya Nutheti , Yifan Peng , Akhil Raj , Ramin Zabih , Curtis L. Cole

The digitization of healthcare has facilitated the sharing and re-using of medical data but has also raised concerns about confidentiality and privacy. HIPAA (Health Insurance Portability and Accountability Act) mandates removing…

Automated deidentification of clinical text data is crucial due to the high cost of manual deidentification, which has been a barrier to sharing clinical text and the advancement of clinical natural language processing. However, creating…

Computation and Language · Computer Science 2023-11-07 Callandra Moore , Jonathan Ranisau , Walter Nelson , Jeremy Petch , Alistair Johnson

Medical imaging has significantly advanced computer-aided diagnosis, yet its re-identification (ReID) risks raise critical privacy concerns, calling for de-identification (DeID) techniques. Unfortunately, existing DeID methods neither…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Yuan Tian , Shuo Wang , Rongzhao Zhang , Zijian Chen , Yankai Jiang , Chunyi Li , Xiangyang Zhu , Fang Yan , Qiang Hu , XiaoSong Wang , Guangtao Zhai

Recent research advances achieve human-level accuracy for de-identifying free-text clinical notes on research datasets, but gaps remain in reproducing this in large real-world settings. This paper summarizes lessons learned from building a…

Computation and Language · Computer Science 2023-12-15 Veysel Kocaman , Hasham Ul Haq , David Talby

Access to medical imaging and associated text data has the potential to drive major advances in healthcare research and patient outcomes. However, the presence of Protected Health Information (PHI) and Personally Identifiable Information…

Machine Learning · Statistics 2025-08-01 Kyle Naddeo , Nikolas Koutsoubis , Rahul Krish , Ghulam Rasool , Nidhal Bouaynaya , Tony OSullivan , Raj Krish

Exploiting natural language processing in the clinical domain requires de-identification, i.e., anonymization of personal information in texts. However, current research considers de-identification and downstream tasks, such as concept…

Computation and Language · Computer Science 2020-05-20 Lukas Lange , Heike Adel , Jannik Strötgen

Unstructured textual data are at the heart of health systems: liaison letters between doctors, operating reports, coding of procedures according to the ICD-10 standard, etc. The details included in these documents make it possible to get to…

Cryptography and Security · Computer Science 2023-10-09 Yakini Tchouka , Jean-François Couchot , Maxime Coulmeau , David Laiymani , Philippe Selles , Azzedine Rahmani

Medical health records and clinical summaries contain a vast amount of important information in textual form that can help advancing research on treatments, drugs and public health. However, the majority of these information is not shared…

Computation and Language · Computer Science 2020-10-13 Nikola Milosevic , Gangamma Kalappa , Hesam Dadafarin , Mahmoud Azimaee , Goran Nenadic

De-identification is the task of detecting protected health information (PHI) in medical text. It is a critical step in sanitizing electronic health records (EHRs) to be shared for research. Automatic de-identification classifierscan…

Computation and Language · Computer Science 2019-06-13 Max Friedrich , Arne Köhn , Gregor Wiedemann , Chris Biemann

De-identification of clinical text remains essential for secondary use of electronic health records (EHRs), yet public benchmarks such as i2b2 2006/2014 are over a decade old and lack the semantic and demographic diversity of modern…

Computation and Language · Computer Science 2026-05-06 Jose D. Posada , David Love , Somalee Datta , Priya Desai

Supporting public health research and the public's situational awareness during a pandemic requires continuous dissemination of infectious disease surveillance data. Legislation, such as the Health Insurance Portability and Accountability…

Cryptography and Security · Computer Science 2022-02-28 J. Thomas Brown , Chao Yan , Weiyi Xia , Zhijun Yin , Zhiyu Wan , Aris Gkoulalas-Divanis , Murat Kantarcioglu , Bradley A. Malin

Ensuring clinical data privacy while preserving utility is critical for AI-driven healthcare and data analytics. Existing de-identification (De-ID) methods, including rule-based techniques, deep learning models, and large language models…

Artificial Intelligence · Computer Science 2025-07-28 Praphul Singh , Charlotte Dzialo , Jangwon Kim , Sumana Srivatsa , Irfan Bulu , Sri Gadde , Krishnaram Kenthapadi

Ensuring the de-identification of medical imaging data is a critical step in enabling safe data sharing. This paper presents a hybrid de-identification framework designed to process Digital Imaging and Communications in Medicine (DICOM)…

Cryptography and Security · Computer Science 2025-09-03 Hamideh Haghiri , Rajesh Baidya , Stefan Dvoretskii , Klaus H. Maier-Hein , Marco Nolden

De-identification is the task of identifying protected health information (PHI) in the clinical text. Existing neural de-identification models often fail to generalize to a new dataset. We propose a simple yet effective data augmentation…

Computation and Language · Computer Science 2020-10-13 Xiang Yue , Shuang Zhou

Duplicate records pose significant challenges in customer relationship management (CRM)and healthcare, often leading to inaccuracies in analytics, impaired user experiences, and compliance risks. Traditional deduplication methods rely…

Machine Learning · Computer Science 2026-03-27 Mohammed Omer Shakeel Ahmed

In this work, we propose a novel problem formulation for de-identification of unstructured clinical text. We formulate the de-identification problem as a sequence to sequence learning problem instead of a token classification problem. Our…

Computation and Language · Computer Science 2021-09-13 Md Monowar Anjum , Noman Mohammed , Xiaoqian Jiang
‹ Prev 1 2 3 10 Next ›