中文
相关论文

相关论文: A primer on synthetic health data

200 篇论文

Data quality is the key factor for the development of trustworthy AI in healthcare. A large volume of curated datasets with controlled confounding factors can help improve the accuracy, robustness and privacy of downstream AI algorithms.…

机器学习 · 计算机科学 2022-09-21 Xiaodan Xing , Huanjun Wu , Lichao Wang , Iain Stenson , May Yong , Javier Del Ser , Simon Walsh , Guang Yang

The widespread use of big data across sectors has raised major privacy concerns, especially when sensitive information is shared or analyzed. Regulations such as GDPR and HIPAA impose strict controls on data handling, making it difficult to…

机器学习 · 计算机科学 2025-12-10 Anantaa Kotal , Anupam Joshi

Differential privacy allows quantifying privacy loss resulting from accessing sensitive personal data. Repeated accesses to underlying data incur increasing loss. Releasing data as privacy-preserving synthetic data would avoid this…

机器学习 · 统计学 2021-06-10 Joonas Jälkö , Eemil Lagerspetz , Jari Haukka , Sasu Tarkoma , Antti Honkela , Samuel Kaski

Institutions collect massive learning traces but they may not disclose it for privacy issues. Synthetic data generation opens new opportunities for research in education. In this paper we present a generative model for educational data that…

计算机与社会 · 计算机科学 2022-07-09 Jill-Jênn Vie , Tomas Rigaux , Sein Minn

For sharing privacy-sensitive data, de-identification is commonly regarded as adequate for safeguarding privacy. Synthetic data is also being considered as a privacy-preserving alternative. Recent successes with numerical and tabular data…

计算与语言 · 计算机科学 2025-03-05 Atiquer Rahman Sarkar , Yao-Shun Chuang , Noman Mohammed , Xiaoqian Jiang

Machine learning systems require representations of the real world for training and testing - they require data, and lots of it. Collecting data at scale has logistical and ethical challenges, and synthetic data promises a solution to these…

计算机与社会 · 计算机科学 2024-05-06 Cedric Deslandes Whitney , Justin Norman

As more tech companies engage in rigorous economic analyses, we are confronted with a data problem: in-house papers cannot be replicated due to use of sensitive, proprietary, or private data. Readers are left to assume that the obscured…

综合经济学 · 经济学 2020-11-10 Allison Koenecke , Hal Varian

This paper presents a novel approach to simulating electronic health records (EHRs) using diffusion probabilistic models (DPMs). Specifically, we demonstrate the effectiveness of DPMs in synthesising longitudinal EHRs that capture…

机器学习 · 计算机科学 2023-03-23 Nicholas I-Hsien Kuo , Louisa Jorm , Sebastiano Barbieri

Differentially private data generation techniques have become a promising solution to the data privacy challenge -- it enables sharing of data while complying with rigorous privacy guarantees, which is essential for scientific progress in…

密码学与安全 · 计算机科学 2022-11-09 Dingfan Chen , Raouf Kerkouche , Mario Fritz

Testing in production-like test environments is an essential part of quality assurance processes in many industries. Provisioning of such test environments, for information-intensive services, involves setting up databases that are…

软件工程 · 计算机科学 2024-07-09 Razieh Behjati , Erik Arisholm , Chao Tan , Margrethe M. Bedregal

Synthetic data is an increasingly popular tool for training deep learning models, especially in computer vision but also in other areas. In this work, we attempt to provide a comprehensive survey of the various directions in the development…

机器学习 · 计算机科学 2019-09-26 Sergey I. Nikolenko

Background: High-level system testing of applications that use data from e-Government services as input requires test data that is real-life-like but where the privacy of personal information is guaranteed. Applications with such strong…

机器学习 · 计算机科学 2026-02-09 Maj-Annika Tammisto , Faiz Ali Shah , Daniel Rodriguez , Dietmar Pfahl

The generation of privacy-preserving synthetic datasets is a promising avenue for overcoming data scarcity in medical AI research. Post-hoc privacy filtering techniques, designed to remove samples containing personally identifiable…

机器学习 · 计算机科学 2025-10-03 Adil Koeken , Alexander Ziller , Moritz Knolle , Daniel Rueckert

Faced with the challenges of patient confidentiality and scientific reproducibility, research on machine learning for health is turning towards the conception of synthetic medical databases. This article presents a brief overview of…

Deep learning models have demonstrated superior performance in several application problems, such as image classification and speech processing. However, creating a deep learning model using health record data requires addressing certain…

机器学习 · 计算机科学 2021-12-14 Amirsina Torfi , Edward A. Fox , Chandan K. Reddy

To develop public health intervention models using microsimulations, extensive personal information about inhabitants is needed, such as socio-demographic, economic and health figures. Data confidentiality is an essential characteristic of…

应用统计 · 统计学 2022-02-10 M. A. Nicolaie , Koen Fussenich , Caroline Ameling , Hendriek C. Boshuizen

Advances in generative models have transformed the field of synthetic image generation for privacy-preserving data synthesis (PPDS). However, the field lacks a comprehensive survey and comparison of synthetic image generation methods across…

密码学与安全 · 计算机科学 2025-06-27 Yunsung Chung , Yunbei Zhang , Nassir Marrouche , Jihun Hamm

Preservation of private user data is of paramount importance for high Quality of Experience (QoE) and acceptability, particularly with services treating sensitive data, such as IT-based health services. Whereas anonymization techniques were…

机器学习 · 计算机科学 2024-03-04 Navid Ashrafi , Vera Schmitt , Robert P. Spang , Sebastian Möller , Jan-Niklas Voigt-Antons

Synthetic data generation has emerged as a promising approach to address the challenges of using sensitive financial data in machine learning applications. By leveraging generative models, such as Generative Adversarial Networks (GANs) and…

机器学习 · 计算机科学 2025-10-31 James Meldrum , Basem Suleiman , Fethi Rabhi , Muhammad Johan Alibasa

Private synthetic data sharing is preferred as it keeps the distribution and nuances of original data compared to summary statistics. The state-of-the-art methods adopt a select-measure-generate paradigm, but measuring large domain…

密码学与安全 · 计算机科学 2023-10-11 Meifan Zhang , Dihang Deng , Lihua Yin