English
Related papers

Related papers: DeID-GPT: Zero-shot Medical Text De-Identification…

200 papers

Unstructured textual data is at the heart of healthcare systems. For obvious privacy reasons, these documents are not accessible to researchers as long as they contain personally identifiable information. One way to share this data while…

Cryptography and Security · Computer Science 2022-11-03 Yakini Tchouka , Jean-François Couchot , David Laiymani

Medical data employed in research frequently comprises sensitive patient health information (PHI), which is subject to rigorous legal frameworks such as the General Data Protection Regulation (GDPR) or the Health Insurance Portability and…

Image and Video Processing · Electrical Eng. & Systems 2024-10-17 Moritz Rempe , Lukas Heine , Constantin Seibold , Fabian Hörst , Jens Kleesiek

The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sharing while preserving downstream utility. This paper presents DeID-Clinic, a multi-layered…

Computation and Language · Computer Science 2026-05-26 Angel Paul , Dhivin Shaji , Lifeng Han , Warren Del-Pinto , Goran Nenadic , Suzan Verberne

De-identification is the task of detecting protected health information (PHI) in medical text. It is a critical step in sanitizing electronic health records (EHRs) to be shared for research. Automatic de-identification classifierscan…

Computation and Language · Computer Science 2019-06-13 Max Friedrich , Arne Köhn , Gregor Wiedemann , Chris Biemann

Unstructured textual data are at the heart of health systems: liaison letters between doctors, operating reports, coding of procedures according to the ICD-10 standard, etc. The details included in these documents make it possible to get to…

Cryptography and Security · Computer Science 2023-10-09 Yakini Tchouka , Jean-François Couchot , Maxime Coulmeau , David Laiymani , Philippe Selles , Azzedine Rahmani

Recent research advances achieve human-level accuracy for de-identifying free-text clinical notes on research datasets, but gaps remain in reproducing this in large real-world settings. This paper summarizes lessons learned from building a…

Computation and Language · Computer Science 2023-12-15 Veysel Kocaman , Hasham Ul Haq , David Talby

Background: Electronic health records (EHRs) are a valuable resource for data-driven medical research. However, the presence of protected health information (PHI) makes EHRs unsuitable to be shared for research purposes. De-identification,…

Computation and Language · Computer Science 2024-04-11 Aleksandar Kovačević , Bojana Bašaragin , Nikola Milošević , Goran Nenadić

Medical health records and clinical summaries contain a vast amount of important information in textual form that can help advancing research on treatments, drugs and public health. However, the majority of these information is not shared…

Computation and Language · Computer Science 2020-10-13 Nikola Milosevic , Gangamma Kalappa , Hesam Dadafarin , Mahmoud Azimaee , Goran Nenadic

We evaluate the performance of four leading solutions for de-identification of unstructured medical text - Azure Health Data Services, AWS Comprehend Medical, OpenAI GPT-4o, and John Snow Labs - on a ground truth dataset of 48 clinical…

Computation and Language · Computer Science 2025-04-02 Veysel Kocaman , Muhammed Santas , Yigit Gul , Mehmet Butgul , David Talby

De-identification in the healthcare setting is an application of NLP where automated algorithms are used to remove personally identifying information of patients (and, sometimes, providers). With the recent rise of generative large language…

Computation and Language · Computer Science 2025-09-19 Kiana Aghakasiri , Noopur Zambare , JoAnn Thai , Carrie Ye , Mayur Mehta , J. Ross Mitchell , Mohamed Abdalla

Clinical free-text data offers immense potential to improve population health research such as richer phenotyping, symptom tracking, and contextual understanding of patient care. However, these data present significant privacy risks due to…

Improving the accuracy and reliability of medical coding reduces clinician burnout and supports revenue cycle processes, freeing providers to focus more on patient care. However, automating the assignment of ICD-10-CM and CPT codes from…

Zero-shot medical image classification is a critical process in real-world scenarios where we have limited access to all possible diseases or large-scale annotated data. It involves computing similarity scores between a query medical image…

Image and Video Processing · Electrical Eng. & Systems 2023-07-06 Jiaxiang Liu , Tianxiang Hu , Yan Zhang , Xiaotang Gai , Yang Feng , Zuozhu Liu

This study examines integrating EHRs and NLP with large language models (LLMs) to improve healthcare data management and patient care. It focuses on using advanced models to create secure, HIPAA-compliant synthetic patient notes for…

Computation and Language · Computer Science 2025-06-03 Yao-Shun Chuang , Atiquer Rahman Sarkar , Yu-Chun Hsu , Noman Mohammed , Xiaoqian Jiang

Medical imaging has significantly advanced computer-aided diagnosis, yet its re-identification (ReID) risks raise critical privacy concerns, calling for de-identification (DeID) techniques. Unfortunately, existing DeID methods neither…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Yuan Tian , Shuo Wang , Rongzhao Zhang , Zijian Chen , Yankai Jiang , Chunyi Li , Xiangyang Zhu , Fang Yan , Qiang Hu , XiaoSong Wang , Guangtao Zhai

Privacy is a human right that sustains patient-provider trust. Clinical notes capture a patient's private vulnerability and individuality, which are used for care coordination and research. Under HIPAA Safe Harbor, these notes are…

Computers and Society · Computer Science 2026-02-10 Lavender Y. Jiang , Xujin Chris Liu , Kyunghyun Cho , Eric K. Oermann

Use of medical data, also known as electronic health records, in research helps develop and advance medical science. However, protecting patient confidentiality and identity while using medical data for analysis is crucial. Medical data can…

Artificial Intelligence · Computer Science 2018-10-17 Vithya Yogarajan , Michael Mayo , Bernhard Pfahringer

The rise of chronic diseases and pandemics like COVID-19 has emphasized the need for effective patient data processing while ensuring privacy through anonymization and de-identification of protected health information (PHI). Anonymized data…

Computation and Language · Computer Science 2024-12-17 Murat Gunay , Bunyamin Keles , Raife Hizlan

Objective: Patient notes in electronic health records (EHRs) may contain critical information for medical investigations. However, the vast majority of medical investigators can only access de-identified notes, in order to protect the…

Computation and Language · Computer Science 2016-06-14 Franck Dernoncourt , Ji Young Lee , Ozlem Uzuner , Peter Szolovits

De-identification is the process of removing 18 protected health information (PHI) from clinical notes in order for the text to be considered not individually identifiable. Recent advances in natural language processing (NLP) has allowed…

Computation and Language · Computer Science 2018-10-04 Kaung Khin , Philipp Burckhardt , Rema Padman
‹ Prev 1 2 3 10 Next ›