English
Related papers

Related papers: DeIDClinic: A Risk-Aware Pseudonymization Framewor…

200 papers

Sharing sensitive texts for scientific purposes requires appropriate techniques to protect the privacy of patients and healthcare personnel. Anonymizing textual data is particularly challenging due to the presence of diverse unstructured…

Computation and Language · Computer Science 2025-02-20 Ibrahim Baroud , Lisa Raithel , Sebastian Möller , Roland Roller

The objective of this study is to address the critical issue of de-identification of clinical reports in order to allow access to data for research purposes, while ensuring patient privacy. The study highlights the difficulties faced in…

Computation and Language · Computer Science 2023-03-24 Xavier Tannier , Perceval Wajsbürt , Alice Calliger , Basile Dura , Alexandre Mouchet , Martin Hilka , Romain Bey

Objective: Patient notes in electronic health records (EHRs) may contain critical information for medical investigations. However, the vast majority of medical investigators can only access de-identified notes, in order to protect the…

Computation and Language · Computer Science 2016-06-14 Franck Dernoncourt , Ji Young Lee , Ozlem Uzuner , Peter Szolovits

Deidentification seeks to anonymize textual data prior to distribution. Automatic deidentification primarily uses supervised named entity recognition from human-labeled data points. We propose an unsupervised deidentification method that…

Computation and Language · Computer Science 2022-10-24 John X. Morris , Justin T. Chiu , Ramin Zabih , Alexander M. Rush

Data sharing is crucial for open science and reproducible research, but the legal sharing of clinical data requires the removal of protected health information from electronic health records. This process, known as de-identification, is…

Machine Learning · Computer Science 2024-01-04 Yuxin Xiao , Shulammite Lim , Tom Joseph Pollard , Marzyeh Ghassemi

Medical data employed in research frequently comprises sensitive patient health information (PHI), which is subject to rigorous legal frameworks such as the General Data Protection Regulation (GDPR) or the Health Insurance Portability and…

Image and Video Processing · Electrical Eng. & Systems 2024-10-17 Moritz Rempe , Lukas Heine , Constantin Seibold , Fabian Hörst , Jens Kleesiek

Data containing personal information is increasingly used to train, fine-tune, or query Large Language Models (LLMs). Text is typically scrubbed of identifying information prior to use, often with tools such as Microsoft's Presidio or…

Computation and Language · Computer Science 2026-02-16 Nataša Krčo , Zexi Yao , Matthieu Meeus , Yves-Alexandre de Montjoye

The rise of chronic diseases and pandemics like COVID-19 has emphasized the need for effective patient data processing while ensuring privacy through anonymization and de-identification of protected health information (PHI). Anonymized data…

Computation and Language · Computer Science 2024-12-17 Murat Gunay , Bunyamin Keles , Raife Hizlan

The robust development of Electronic Health Records (EHRs) causes a significant growth in sharing EHRs for clinical research. However, such a sharing makes it difficult to protect patient's privacy. A number of automated de-identification…

Cryptography and Security · Computer Science 2012-11-19 Jie Qian , Nafees Qamar

Many models are pretrained on redacted text for privacy reasons. Clinical foundation models are often trained on de-identified text, which uses special syntax (masked) text in place of protected health information. Even though these models…

Computation and Language · Computer Science 2025-06-18 Paul Landes , Aaron J Chaise , Tarak Nath Nandi , Ravi K Madduri

Objective: The use of routinely-acquired medical data for research purposes requires the protection of patient confidentiality via data anonymisation. The objective of this work is to calculate the risk of re-identification arising from a…

Machine Learning · Computer Science 2022-04-01 Anna Antoniou , Giacomo Dossena , Julia MacMillan , Steven Hamblin , David Clifton , Paula Petrone

Background: Electronic health records (EHRs) are a valuable resource for data-driven medical research. However, the presence of protected health information (PHI) makes EHRs unsuitable to be shared for research purposes. De-identification,…

Computation and Language · Computer Science 2024-04-11 Aleksandar Kovačević , Bojana Bašaragin , Nikola Milošević , Goran Nenadić

Face de-identification (DeID) has been widely studied for common scenes, but remains under-researched for medical scenes, mostly due to the lack of large-scale patient face datasets. In this paper, we release MeMa, consisting of over 40,000…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yuan Tian , Shuo Wang , Guangtao Zhai

Objectives; The accumulation and usefulness of clinical data have increased with IT development. While using clinical data that needs to be identifiable to obtain meaningful information, it is essential to ensure that data is de-identified…

Cryptography and Security · Computer Science 2018-04-16 Jipmin Jung , Phillip Park , Jaedong Lee , Hyein Lee , Geonkook Lee , Hyosoung Cha

De-identification is the task of detecting privacy-related entities in text, such as person names, emails and contact data. It has been well-studied within the medical domain. The need for de-identification technology is increasing, as…

Computation and Language · Computer Science 2021-05-25 Kristian Nørgaard Jensen , Mike Zhang , Barbara Plank

To prove that a dataset is sufficiently anonymized, many privacy policies suggest that a re-identification risk assessment be performed, but do not provide a precise methodology for doing so, leaving the industry alone with the problem.…

Cryptography and Security · Computer Science 2025-01-22 Louis-Philippe Sondeck , Maryline Laurent

Documents revealing sensitive information about individuals must typically be de-identified. This de-identification is often done by masking all mentions of personally identifiable information (PII), thereby making it more difficult to…

Computation and Language · Computer Science 2025-05-20 Lucas Georges Gabriel Charpentier , Pierre Lison

Face anonymization aims to conceal the visual identity of a face to safeguard the individual's privacy. Traditional methods like blurring and pixelation can largely remove identifying features, but these techniques significantly degrade…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Lin Yuan , Kai Liang , Xiong Li , Tao Wu , Nannan Wang , Xinbo Gao

For sharing privacy-sensitive data, de-identification is commonly regarded as adequate for safeguarding privacy. Synthetic data is also being considered as a privacy-preserving alternative. Recent successes with numerical and tabular data…

Computation and Language · Computer Science 2025-03-05 Atiquer Rahman Sarkar , Yao-Shun Chuang , Noman Mohammed , Xiaoqian Jiang

Due to the rapid advancement of Large Language Model (LLM), the whole community eagerly consumes any available text data in order to train the LLM. Currently, large portion of the available text data are collected from internet, which has…

Artificial Intelligence · Computer Science 2024-06-21 Ya-Lun Li