中文
相关论文

相关论文: SHIELD: A Diverse Clinical Note Dataset and Distil…

200 篇论文

De-identification is the process of removing 18 protected health information (PHI) from clinical notes in order for the text to be considered not individually identifiable. Recent advances in natural language processing (NLP) has allowed…

计算与语言 · 计算机科学 2018-10-04 Kaung Khin , Philipp Burckhardt , Rema Padman

The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sharing while preserving downstream utility. This paper presents DeID-Clinic, a multi-layered…

计算与语言 · 计算机科学 2026-05-26 Angel Paul , Dhivin Shaji , Lifeng Han , Warren Del-Pinto , Goran Nenadic , Suzan Verberne

Despite the advances in digital healthcare systems offering curated structured knowledge, much of the critical information still lies in large volumes of unlabeled and unstructured clinical texts. These texts, which often contain protected…

We present SHIELD, a novel methodology for automated and integrated safety signal detection in clinical trials. SHIELD combines disproportionality analysis with semantic clustering of adverse event (AE) terms applied to MedDRA term…

计算与语言 · 计算机科学 2026-02-24 Francois Vandenhende , Anna Georgiou , Theodoros Psaras , Ellie Karekla

Protected health information (PHI) de-identification is critical for enabling the safe reuse of clinical notes, yet evaluating and comparing PHI de-identification models typically depends on costly, small-scale expert annotations. We…

人工智能 · 计算机科学 2025-11-19 Guanchen Wu , Zuhui Chen , Yuzhang Xie , Carl Yang

Background: Electronic health records (EHRs) are a valuable resource for data-driven medical research. However, the presence of protected health information (PHI) makes EHRs unsuitable to be shared for research purposes. De-identification,…

计算与语言 · 计算机科学 2024-04-11 Aleksandar Kovačević , Bojana Bašaragin , Nikola Milošević , Goran Nenadić

Large language models (LLMs) excel at clinical information extraction but their computational demands limit practical deployment. Knowledge distillation--the process of transferring knowledge from larger to smaller models--offers a…

计算与语言 · 计算机科学 2025-01-03 Karthik S. Vedula , Annika Gupta , Akshay Swaminathan , Ivan Lopez , Suhana Bedi , Nigam H. Shah

Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporary deidentification…

计算与语言 · 计算机科学 2024-10-23 John X. Morris , Thomas R. Campion , Sri Laasya Nutheti , Yifan Peng , Akhil Raj , Ramin Zabih , Curtis L. Cole

The rise of chronic diseases and pandemics like COVID-19 has emphasized the need for effective patient data processing while ensuring privacy through anonymization and de-identification of protected health information (PHI). Anonymized data…

计算与语言 · 计算机科学 2024-12-17 Murat Gunay , Bunyamin Keles , Raife Hizlan

This study examines integrating EHRs and NLP with large language models (LLMs) to improve healthcare data management and patient care. It focuses on using advanced models to create secure, HIPAA-compliant synthetic patient notes for…

计算与语言 · 计算机科学 2025-06-03 Yao-Shun Chuang , Atiquer Rahman Sarkar , Yu-Chun Hsu , Noman Mohammed , Xiaoqian Jiang

Recent research advances achieve human-level accuracy for de-identifying free-text clinical notes on research datasets, but gaps remain in reproducing this in large real-world settings. This paper summarizes lessons learned from building a…

计算与语言 · 计算机科学 2023-12-15 Veysel Kocaman , Hasham Ul Haq , David Talby

Identifying medication discontinuations in electronic health records (EHRs) is vital for patient safety but is often hindered by information being buried in unstructured notes. This study aims to evaluate the capabilities of advanced…

计算与语言 · 计算机科学 2025-11-10 Chong Shao , Douglas Snyder , Chiran Li , Bowen Gu , Kerry Ngan , Chun-Ting Yang , Jiageng Wu , Richard Wyss , Kueiyu Joshua Lin , Jie Yang

Removing Personally Identifiable Information (PII) from clinical notes in Electronic Health Records (EHRs) is essential for research and AI development. While Large Language Models (LLMs) are powerful, their high computational costs and the…

计算与语言 · 计算机科学 2025-10-23 Prakrithi Shivaprakash , Lekhansh Shukla , Animesh Mukherjee , Prabhat Chand , Pratima Murthy

Automated labeling of chest X-ray reports is essential for enabling downstream tasks such as training image-based diagnostic models, population health studies, and clinical decision support. However, the high variability, complexity, and…

计算与语言 · 计算机科学 2025-05-06 Brian Wong , Kaito Tanaka

Host-based intrusion detection system (HIDS) is a key defense component to protect the organizations from advanced threats like Advanced Persistent Threats (APT). By analyzing the fine-grained logs with approaches like data provenance, HIDS…

密码学与安全 · 计算机科学 2025-07-16 Danyu Sun , Jinghuai Zhang , Jiacen Xu , Yu Zheng , Yuan Tian , Zhou Li

Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full retraining. Most existing methods require a retain set of…

机器学习 · 计算机科学 2026-05-11 Zizhao Hu , Ameya Godbole , Johnny Tian-Zheng Wei , Mohammad Rostami , Jesse Thomason , Robin Jia

De-identification is the task of detecting protected health information (PHI) in medical text. It is a critical step in sanitizing electronic health records (EHRs) to be shared for research. Automatic de-identification classifierscan…

计算与语言 · 计算机科学 2019-06-13 Max Friedrich , Arne Köhn , Gregor Wiedemann , Chris Biemann

Electronic Health Records (EHRs) have become the primary form of medical data-keeping across the United States. Federal law restricts the sharing of any EHR data that contains protected health information (PHI). De-identification, the…

计算与语言 · 计算机科学 2021-03-26 Abdullah Ahmed , Adeel Abbasi , Carsten Eickhoff

Free-text clinical notes detail all aspects of patient care and have great potential to facilitate quality improvement and assurance initiatives as well as advance clinical research. However, concerns about patient privacy and…

计算与语言 · 计算机科学 2021-02-23 Nicholas Dobbins , David Wayne , Kahyun Lee , Özlem Uzuner , Meliha Yetisgen

We introduce Secure Haplotype Imputation Employing Local Differential privacy (SHIELD), a program for accurately estimating the genotype of target samples at markers that are not directly assayed by array-based genotyping platforms while…

定量方法 · 定量生物学 2023-09-15 Marc Harary
‹ 上一页 1 2 3 10 下一页 ›