中文
相关论文

相关论文: BRATsynthetic: Text De-identification using a Mark…

200 篇论文

Objective: To enhance automated de-identification of radiology reports by scaling transformer-based models through extensive training datasets and benchmarking performance against commercial cloud vendor systems for protected health…

In the field of machine learning, domain-specific annotated data is an invaluable resource for training effective models. However, in the medical domain, this data often includes Personal Health Information (PHI), raising significant…

计算与语言 · 计算机科学 2024-09-13 Tal Baumel , Andre Manoel , Daniel Jones , Shize Su , Huseyin Inan , Aaron , Bornstein , Robert Sim

This study examines integrating EHRs and NLP with large language models (LLMs) to improve healthcare data management and patient care. It focuses on using advanced models to create secure, HIPAA-compliant synthetic patient notes for…

计算与语言 · 计算机科学 2025-06-03 Yao-Shun Chuang , Atiquer Rahman Sarkar , Yu-Chun Hsu , Noman Mohammed , Xiaoqian Jiang

Free-text clinical notes detail all aspects of patient care and have great potential to facilitate quality improvement and assurance initiatives as well as advance clinical research. However, concerns about patient privacy and…

计算与语言 · 计算机科学 2021-02-23 Nicholas Dobbins , David Wayne , Kahyun Lee , Özlem Uzuner , Meliha Yetisgen

This paper explores the strategic use of modern synthetic data generation and advanced data perturbation techniques to enhance security, maintain analytical utility, and improve operational efficiency when managing large datasets, with a…

密码学与安全 · 计算机科学 2025-04-29 Anantha Sharma , Swetha Devabhaktuni , Eklove Mohan

Objective: Patient notes in electronic health records (EHRs) may contain critical information for medical investigations. However, the vast majority of medical investigators can only access de-identified notes, in order to protect the…

计算与语言 · 计算机科学 2016-06-14 Franck Dernoncourt , Ji Young Lee , Ozlem Uzuner , Peter Szolovits

Patient notes contain a wealth of information of potentially great interest to medical investigators. However, to protect patients' privacy, Protected Health Information (PHI) must be removed from the patient notes before they can be…

计算与语言 · 计算机科学 2016-11-01 Ji Young Lee , Franck Dernoncourt , Ozlem Uzuner , Peter Szolovits

De-identification is the task of detecting protected health information (PHI) in medical text. It is a critical step in sanitizing electronic health records (EHRs) to be shared for research. Automatic de-identification classifierscan…

计算与语言 · 计算机科学 2019-06-13 Max Friedrich , Arne Köhn , Gregor Wiedemann , Chris Biemann

In decision-making problems, the actions of an agent may reveal sensitive information that drives its decisions. For instance, a corporation's investment decisions may reveal its sensitive knowledge about market dynamics. To prevent this…

系统与控制 · 电气工程与系统科学 2020-04-17 Parham Gohari , Matthew Hale , Ufuk Topcu

Electronic health records (EHR) often contain sensitive medical information about individual patients, posing significant limitations to sharing or releasing EHR data for downstream learning and inferential tasks. We use normalizing flows…

机器学习 · 统计学 2023-02-14 Bingyue Su , Yu Wang , Daniele E. Schiavazzi , Fang Liu

Access to medical imaging and associated text data has the potential to drive major advances in healthcare research and patient outcomes. However, the presence of Protected Health Information (PHI) and Personally Identifiable Information…

For sharing privacy-sensitive data, de-identification is commonly regarded as adequate for safeguarding privacy. Synthetic data is also being considered as a privacy-preserving alternative. Recent successes with numerical and tabular data…

计算与语言 · 计算机科学 2025-03-05 Atiquer Rahman Sarkar , Yao-Shun Chuang , Noman Mohammed , Xiaoqian Jiang

Publishing trajectory data (individual's movement information) is very useful, but it also raises privacy concerns. To handle the privacy concern, in this paper, we apply differential privacy, the standard technique for data privacy,…

密码学与安全 · 计算机科学 2022-10-06 Haiming Wang , Zhikun Zhang , Tianhao Wang , Shibo He , Michael Backes , Jiming Chen , Yang Zhang

The generation of privacy-preserving synthetic datasets is a promising avenue for overcoming data scarcity in medical AI research. Post-hoc privacy filtering techniques, designed to remove samples containing personally identifiable…

机器学习 · 计算机科学 2025-10-03 Adil Koeken , Alexander Ziller , Moritz Knolle , Daniel Rueckert

We propose a method for the release of differentially private synthetic datasets. In many contexts, data contain sensitive values which cannot be released in their original form in order to protect individuals' privacy. Synthetic data is a…

统计方法学 · 统计学 2018-05-25 Joshua Snoke , Aleksandra Slavković

Electronic Health Records (EHRs) are rich sources of patient-level data, offering valuable resources for medical data analysis. However, privacy concerns often restrict access to EHRs, hindering downstream analysis. Current EHR…

机器学习 · 计算机科学 2024-12-03 Muhang Tian , Bernie Chen , Allan Guo , Shiyi Jiang , Anru R. Zhang

As more and more personal photos are shared and tagged in social media, avoiding privacy risks such as unintended recognition becomes increasingly challenging. We propose a new hybrid approach to obfuscate identities in photos by head…

计算机视觉与模式识别 · 计算机科学 2018-07-25 Qianru Sun , Ayush Tewari , Weipeng Xu , Mario Fritz , Christian Theobalt , Bernt Schiele

Objective: To comparatively evaluate several transformer model architectures at identifying protected health information (PHI) in the i2b2/UTHealth 2014 clinical text de-identification challenge corpus. Methods: The i2b2/UTHealth 2014…

计算与语言 · 计算机科学 2022-04-15 Christopher Meaney , Wali Hakimpour , Sumeet Kalia , Rahim Moineddin

Veterinary electronic health records (vEHRs) contain privacy-sensitive identifiers that limit secondary use. While PetEVAL provides a benchmark for veterinary de-identification, the domain remains low-resource. This study evaluates whether…

密码学与安全 · 计算机科学 2026-01-16 David Brundage

Single nucleotide polymorphism (SNP) datasets are fundamental to genetic studies but pose significant privacy risks when shared. The correlation of SNPs with each other makes strong adversarial attacks such as masked-value reconstruction,…

机器学习 · 计算机科学 2025-10-08 Shadi Rahimian , Mario Fritz
‹ 上一页 1 2 3 10 下一页 ›