中文
相关论文

相关论文: A simplified approach to generating synthetic data…

200 篇论文

We report on our experiences of helping staff of the Scottish Longitudinal Study to create synthetic extracts that can be released to users. In particular, we focus on how the synthesis process can be tailored to produce synthetic extracts…

应用统计 · 统计学 2017-12-13 Gillian M. Raab , Beata Nowok , Chris Dibben

A common approach to synthetic data is to sample from a fitted model. We show that under general assumptions, this approach results in a sample with inefficient estimators and whose joint distribution is inconsistent with the true…

统计理论 · 数学 2026-02-18 Jordan Awan , Zhanrui Cai

Access to individual-level health data is essential for gaining new insights and advancing science. In particular, modern methods based on artificial intelligence rely on the availability of and access to large datasets. In the health…

Deep neural networks have become prevalent in human analysis, boosting the performance of applications, such as biometric recognition, action recognition, as well as person re-identification. However, the performance of such networks scales…

计算机视觉与模式识别 · 计算机科学 2022-08-22 Indu Joshi , Marcel Grimmer , Christian Rathgeb , Christoph Busch , Francois Bremond , Antitza Dantcheva

Synthetic data has gained significant momentum thanks to sophisticated machine learning tools that enable the synthesis of high-dimensional datasets. However, many generation techniques do not give the data controller control over what…

The rapid growth in data availability has facilitated research and development, yet not all industries have benefited equally due to legal and privacy constraints. The healthcare sector faces significant challenges in utilizing patient data…

统计方法学 · 统计学 2025-11-18 Katariina Perkonoja , Kari Auranen , Joni Virta

The dissemination of synthetic data can be an effective means of making information from sensitive data publicly available while reducing the risk of disclosure associated with releasing the sensitive data directly. While mechanisms exist…

统计方法学 · 统计学 2021-09-23 Harrison Quick

Introduction: The amount of data generated by original research is growing exponentially. Publicly releasing them is recommended to comply with the Open Science principles. However, data collected from human participants cannot be released…

机器学习 · 统计学 2023-10-11 Rémy Chapelle , Bruno Falissard

The generation of synthetic data is an essential tool to study complex systems, allowing for example to test models of these in precisely controlled settings, or to parametrize simulation models when data is missing. This paper focuses on…

应用统计 · 统计学 2019-11-25 Juste Raimbault

The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative to real-world data…

计算与语言 · 计算机科学 2025-07-25 Tevin Atwal , Chan Nam Tieu , Yefeng Yuan , Zhan Shi , Yuhong Liu , Liang Cheng

Motivated by privacy concerns in long-term longitudinal studies in medical and social science research, we study the problem of continually releasing differentially private synthetic data from longitudinal data collections. We introduce a…

数据结构与算法 · 计算机科学 2024-05-28 Mark Bun , Marco Gaboardi , Marcel Neunhoeffer , Wanrong Zhang

Recent advances in generative modelling have led many to see synthetic data as the go-to solution for a range of problems around data access, scarcity, and under-representation. In this paper, we study three prominent use cases: (1) Sharing…

机器学习 · 计算机科学 2026-02-04 Bogdan Kulynych , Theresa Stadler , Jean Louis Raisaro , Carmela Troncoso

This paper introduces a new latent variable generative model able to handle high dimensional longitudinal data and relying on variational inference. The time dependency between the observations of an input sequence is modelled using…

机器学习 · 统计学 2023-03-28 Clément Chadebec , Stéphanie Allassonnière

Synthetic data generation is a powerful tool for privacy protection when considering public release of record-level data files. Initially proposed about three decades ago, it has generated significant research and application interest. To…

统计方法学 · 统计学 2023-08-03 Jingchen Hu , Claire McKay Bowen

Generating synthetic data, with or without differential privacy, has attracted significant attention as a potential solution to the dilemma between making data easily available, and the privacy of data subjects. Several works have shown…

统计方法学 · 统计学 2023-11-01 Ossi Räisä , Joonas Jälkö , Antti Honkela

This paper demonstrates the potential of statistical disclosure control for protecting the data used to train recommender systems. Specifically, we use a synthetic data generation approach to hide specific information in the user-item…

信息检索 · 计算机科学 2020-08-11 Manel Slokom , Martha Larson , Alan Hanjalic

The US Decennial Census provides valuable data for both research and policy purposes. Census data are subject to a variety of disclosure avoidance techniques prior to release in order to preserve respondent confidentiality. While many are…

计算机与社会 · 计算机科学 2025-10-02 Cynthia Dwork , Kristjan Greenewald , Manish Raghavan

When seeking to release public use files for confidential data, statistical agencies can generate fully synthetic data. We propose an approach for making fully synthetic data from surveys collected with complex sampling designs. Our…

统计方法学 · 统计学 2024-04-30 Shirley Mathur , Yajuan Si , Jerome P. Reiter

Safe and reliable disclosure of information from confidential data is a challenging statistical problem. A common approach considers the generation of synthetic data, to be disclosed instead of the original data. Efficient approaches ought…

统计方法学 · 统计学 2024-03-04 Larissa N. A. Martins , Flávio B. Gonçalves , Thais P. Galletti

Machine learning heavily relies on data, but real-world applications often encounter various data-related issues. These include data of poor quality, insufficient data points leading to under-fitting of machine learning models, and…

‹ 上一页 1 2 3 10 下一页 ›