English
Related papers

Related papers: PHICON: Improving Generalization of Clinical Text …

200 papers

Many models are pretrained on redacted text for privacy reasons. Clinical foundation models are often trained on de-identified text, which uses special syntax (masked) text in place of protected health information. Even though these models…

Computation and Language · Computer Science 2025-06-18 Paul Landes , Aaron J Chaise , Tarak Nath Nandi , Ravi K Madduri

De-identification is the task of detecting protected health information (PHI) in medical text. It is a critical step in sanitizing electronic health records (EHRs) to be shared for research. Automatic de-identification classifierscan…

Computation and Language · Computer Science 2019-06-13 Max Friedrich , Arne Köhn , Gregor Wiedemann , Chris Biemann

Free-text clinical notes detail all aspects of patient care and have great potential to facilitate quality improvement and assurance initiatives as well as advance clinical research. However, concerns about patient privacy and…

Computation and Language · Computer Science 2021-02-23 Nicholas Dobbins , David Wayne , Kahyun Lee , Özlem Uzuner , Meliha Yetisgen

Objective: To enhance automated de-identification of radiology reports by scaling transformer-based models through extensive training datasets and benchmarking performance against commercial cloud vendor systems for protected health…

Computation and Language · Computer Science 2025-11-24 Eva Prakash , Maayane Attias , Pierre Chambon , Justin Xu , Steven Truong , Jean-Benoit Delbrouck , Tessa Cook , Curtis Langlotz

The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sharing while preserving downstream utility. This paper presents DeID-Clinic, a multi-layered…

Computation and Language · Computer Science 2026-05-26 Angel Paul , Dhivin Shaji , Lifeng Han , Warren Del-Pinto , Goran Nenadic , Suzan Verberne

Medical data employed in research frequently comprises sensitive patient health information (PHI), which is subject to rigorous legal frameworks such as the General Data Protection Regulation (GDPR) or the Health Insurance Portability and…

Image and Video Processing · Electrical Eng. & Systems 2024-10-17 Moritz Rempe , Lukas Heine , Constantin Seibold , Fabian Hörst , Jens Kleesiek

Exploiting natural language processing in the clinical domain requires de-identification, i.e., anonymization of personal information in texts. However, current research considers de-identification and downstream tasks, such as concept…

Computation and Language · Computer Science 2020-05-20 Lukas Lange , Heike Adel , Jannik Strötgen

Text de-identification techniques are often used to mask personally identifiable information (PII) from documents. Their ability to conceal the identity of the individuals mentioned in a text is, however, hard to measure. Recent work has…

Computation and Language · Computer Science 2025-10-13 Lucas Georges Gabriel Charpentier , Pierre Lison

Patient notes contain a wealth of information of potentially great interest to medical investigators. However, to protect patients' privacy, Protected Health Information (PHI) must be removed from the patient notes before they can be…

Computation and Language · Computer Science 2016-11-01 Ji Young Lee , Franck Dernoncourt , Ozlem Uzuner , Peter Szolovits

Access to medical imaging and associated text data has the potential to drive major advances in healthcare research and patient outcomes. However, the presence of Protected Health Information (PHI) and Personally Identifiable Information…

Machine Learning · Statistics 2025-08-01 Kyle Naddeo , Nikolas Koutsoubis , Rahul Krish , Ghulam Rasool , Nidhal Bouaynaya , Tony OSullivan , Raj Krish

Abbreviation disambiguation is important for automated clinical note processing due to the frequent use of abbreviations in clinical settings. Current models for automated abbreviation disambiguation are restricted by the scarcity and…

Machine Learning · Computer Science 2019-12-16 Marta Skreta , Aryan Arbabi , Jixuan Wang , Michael Brudno

Clinical notes often describe the most important aspects of a patient's physiology and are therefore critical to medical research. However, these notes are typically inaccessible to researchers without prior removal of sensitive protected…

Computation and Language · Computer Science 2018-03-08 Willie Boag , Tristan Naumann , Peter Szolovits

This research on data extraction methods applies recent advances in natural language processing to evidence synthesis based on medical texts. Texts of interest include abstracts of clinical trials in English and in multilingual contexts.…

Computation and Language · Computer Science 2020-01-31 Lena Schmidt , Julie Weeds , Julian P. T. Higgins

Removing patient-specific information from medical images is crucial to enable sharing and open science without compromising patient identities. However, many methods currently used for deidentification have negative effects on downstream…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Adrienne Kline , Abhijit Gaonkar , Daniel Pittman , Chris Kuehn , Nils Forkert

Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporary deidentification…

Computation and Language · Computer Science 2024-10-23 John X. Morris , Thomas R. Campion , Sri Laasya Nutheti , Yifan Peng , Akhil Raj , Ramin Zabih , Curtis L. Cole

De-identification of data used for automatic speech recognition modeling is a critical component in protecting privacy, especially in the medical domain. However, simply removing all personally identifiable information (PII) from end-to-end…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-13 Martin Flechl , Shou-Chun Yin , Junho Park , Peter Skala

Deep learning models with large learning capacities often overfit to medical imaging datasets. This is because training sets are often relatively small due to the significant time and financial costs incurred in medical data acquisition and…

Computer Vision and Pattern Recognition · Computer Science 2021-03-16 Lok Hin Lee , Yuan Gao , J. Alison Noble

De-identification is the process of removing 18 protected health information (PHI) from clinical notes in order for the text to be considered not individually identifiable. Recent advances in natural language processing (NLP) has allowed…

Computation and Language · Computer Science 2018-10-04 Kaung Khin , Philipp Burckhardt , Rema Padman

In this work, we propose a novel problem formulation for de-identification of unstructured clinical text. We formulate the de-identification problem as a sequence to sequence learning problem instead of a token classification problem. Our…

Computation and Language · Computer Science 2021-09-13 Md Monowar Anjum , Noman Mohammed , Xiaoqian Jiang

Protected health information (PHI) de-identification is critical for enabling the safe reuse of clinical notes, yet evaluating and comparing PHI de-identification models typically depends on costly, small-scale expert annotations. We…

Artificial Intelligence · Computer Science 2025-11-19 Guanchen Wu , Zuhui Chen , Yuzhang Xie , Carl Yang
‹ Prev 1 2 3 10 Next ›