中文
相关论文

相关论文: Performance of Automatic De-identification Across …

200 篇论文

Federated Learning (FL) offers a promising approach for training clinical AI models without centralizing sensitive patient data. However, its real-world adoption is hindered by challenges related to privacy, resource constraints, and…

Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stages because of nonspecific clinical presentations. However, to our knowledge, there is no…

计算与语言 · 计算机科学 2026-05-28 An Dao , Nhan Ly , Thao Tran , Yuji Matsumoto , Akiko Aizawa

A crucial step within secondary analysis of electronic health records (EHRs) is to identify the patient cohort under investigation. While EHRs contain medical billing codes that aim to represent the conditions and treatments patients may…

Health literacy is a critical determinant of patient outcomes, yet current screening tools are not always feasible and differ considerably in the number of items, question format, and dimensions of health literacy they capture, making…

Although data-driven methods usually have noticeable performance on disease diagnosis and treatment, they are suspected of leakage of privacy due to collecting data for model training. Recently, federated learning provides a secure and…

人工智能 · 计算机科学 2023-06-27 Yawei Zhao , Qinghe Liu , Xinwang Liu , Kunlun He

Data sharing for medical research has been difficult as open-sourcing clinical data may violate patient privacy. Traditional methods for face de-identification wipe out facial information entirely, making it impossible to analyze facial…

计算机视觉与模式识别 · 计算机科学 2020-03-03 Bingquan Zhu , Hao Fang , Yanan Sui , Luming Li

Protecting patient data privacy is a critical concern when deploying machine learning algorithms in healthcare. Differential privacy (DP) is a common method for preserving privacy in such settings and, in this work, we examine two key…

机器学习 · 计算机科学 2024-12-10 Ali Dadsetan , Dorsa Soleymani , Xijie Zeng , Frank Rudzicz

Anonymized data is highly valuable to both businesses and researchers. A large body of research has however shown the strong limits of the de-identification release-and-forget model, where data is anonymized and shared. This has led to the…

密码学与安全 · 计算机科学 2019-10-31 Andrea Gadotti , Florimond Houssiau , Luc Rocher , Benjamin Livshits , Yves-Alexandre de Montjoye

Physician burnout in the United States has reached critical levels, driven in part by the administrative burden of Electronic Health Record (EHR) documentation and complex diagnostic codes. To relieve this strain and maintain strict patient…

信息检索 · 计算机科学 2026-03-25 Peter Hartnett , Chung-Chi Huang , Sarah Hartnett , David Hartnett

Balancing privacy and predictive utility remains a central challenge for machine learning in healthcare. In this paper, we develop Syfer, a neural obfuscation method to protect against re-identification attacks. Syfer composes trained…

Background: The increasing use of artificial intelligence (AI) in healthcare documentation necessitates robust methods for evaluating the quality of AI-generated medical notes compared to those written by humans. This paper introduces an…

人机交互 · 计算机科学 2025-03-24 Iyad Sultan

Clinician notes are a rich source of patient information but often contain inconsistencies due to varied writing styles, colloquialisms, abbreviations, medical jargon, grammatical errors, and non-standard formatting. These inconsistencies…

计算与语言 · 计算机科学 2025-01-03 Daniel B. Hier , Michael D. Carrithers , Thanh Son Do , Tayo Obafemi-Ajayi

Leveraging knowledge from electronic health records (EHRs) to predict a patient's condition is essential to the effective delivery of appropriate care. Clinical notes of patient EHRs contain valuable information from healthcare…

计算与语言 · 计算机科学 2023-05-18 Nayeon Kim , Yinhua Piao , Sun Kim

Documents revealing sensitive information about individuals must typically be de-identified. This de-identification is often done by masking all mentions of personally identifiable information (PII), thereby making it more difficult to…

计算与语言 · 计算机科学 2025-05-20 Lucas Georges Gabriel Charpentier , Pierre Lison

Dementia is under-recognized in the community, under-diagnosed by healthcare professionals, and under-coded in claims data. Information on cognitive dysfunction, however, is often found in unstructured clinician notes within medical records…

Patient-controlled data-sharing systems are increasingly promoted as a way to empower patients with greater autonomy over their health data. Yet it remains unclear how different stakeholders, especially patients and health system leaders,…

We curated WikiPII, an automatically labeled dataset composed of Wikipedia biography pages, annotated for personal information extraction. Although automatic annotation can lead to a high degree of label noise, it is an inexpensive process…

计算与语言 · 计算机科学 2021-05-20 Rajitha Hathurusinghe , Isar Nejadgholi , Miodrag Bolic

Researchers increasingly use data on social and economic networks to study a range of social science questions, but releasing statistics derived from networks can raise significant privacy concerns. We show how to release network…

应用统计 · 统计学 2026-03-17 Tom A. Rutter , Yuxin Liu , M. Amin Rahimian

Large Transformers pretrained over clinical notes from Electronic Health Records (EHR) have afforded substantial gains in performance on predictive clinical tasks. The cost of training such models (and the necessity of data access to do so)…

计算与语言 · 计算机科学 2021-04-26 Eric Lehman , Sarthak Jain , Karl Pichotta , Yoav Goldberg , Byron C. Wallace

Removing personally identifiable information (PII) from texts is necessary to comply with various data protection regulations and to enable data sharing without compromising privacy. However, recent works show that documents sanitized by…

计算与语言 · 计算机科学 2026-03-16 Sebastian Ochs , Ivan Habernal