中文
相关论文

相关论文: Fully Synthetic Data for Complex Surveys

200 篇论文

Population censuses are vital to public policy decision-making. They provide insight into human resources, demography, culture, and economic structure at local, regional, and national levels. However, such surveys are very expensive…

机器学习 · 计算机科学 2024-05-17 Bhavesh Neekhra , Kshitij Kapoor , Debayan Gupta

In general, to draw robust conclusions from a dataset, all the analyzed population must be represented on said dataset. Having a dataset that does not fulfill this condition normally leads to selection bias. Additionally, graphs have been…

机器学习 · 计算机科学 2022-05-30 Axel Wassington , Sergi Abadal

There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in…

统计方法学 · 统计学 2023-04-11 Tian Gu , Jeremy M. G. Taylor , Bhramar Mukherjee

Synthetic data generation offers promise for addressing data scarcity and privacy concerns in educational technology, yet practitioners lack empirical guidance for selecting between traditional resampling techniques and modern deep learning…

机器学习 · 计算机科学 2026-04-24 Tapiwa Amion Chinodakufa , Ashfaq Ali Shafin , Khandaker Mamun Ahmed

Recent studies have highlighted the benefits of generating multiple synthetic datasets for supervised learning, from increased accuracy to more effective model selection and uncertainty estimation. These benefits have clear empirical…

机器学习 · 计算机科学 2025-04-28 Ossi Räisä , Antti Honkela

The performance of supervised deep learning algorithms depends significantly on the scale, quality and diversity of the data used for their training. Collecting and manually annotating large amount of data can be both time-consuming and…

计算机视觉与模式识别 · 计算机科学 2021-07-02 C. Symeonidis , P. Nousi , P. Tosidis , K. Tsampazis , N. Passalis , A. Tefas , N. Nikolaidis

Data privacy has increasingly become a daunting challenge because it limits data availability, which is essential in estimating statistical models such as generalized linear mixed models. Access to personal data often involves considerable…

统计方法学 · 统计学 2026-05-05 Marie Analiz April Limpoco , Christel Faes , Niel Hens

Implementing Bayesian inference is often computationally challenging in applications involving complex models, and sometimes calculating the likelihood itself is difficult. Synthetic likelihood is one approach for carrying out inference…

统计计算 · 统计学 2021-03-15 David T. Frazier , David J. Nott , Christopher Drovandi , Robert Kohn

The comparison of subnational areas is ubiquitous but survey samples of these areas are often biased or prohibitively small. Researchers turn to methods such as multilevel regression and poststratification (MRP) to improve the efficiency of…

统计方法学 · 统计学 2021-05-13 Shiro Kuriwaki , Soichiro Yamauchi

We compare a sample-free method proposed by Gargiulo et al. (2010) and a sample-based method proposed by Ye et al. (2009) for generating a synthetic population, organised in households, from various statistics. We generate a reference…

应用统计 · 统计学 2018-12-27 Maxime Lenormand , Guillaume Deffuant

The U.S. Census Bureau provides an estimate of the true population as a supplement to the basic census numbers. This estimate is constructed from data in a post-censal survey. The overall procedure is referred to as dual system estimation.…

应用统计 · 统计学 2008-12-18 Lawrence Brown , Zhanyun Zhao

Although many AI applications of interest require specialized multi-modal models, relevant data to train such models is inherently scarce or inaccessible. Filling these gaps with human annotators is prohibitively expensive, error-prone, and…

人工智能 · 计算机科学 2026-04-01 Tim R. Davidson , Benoit Seguin , Enrico Bacis , Cesar Ilharco , Hamza Harkous

We consider the problem of synthetically generating data that can closely resemble human decisions made in the context of an interactive human-AI system like a computer game. We propose a novel algorithm that can generate synthetic,…

机器学习 · 计算机科学 2023-04-17 Bryan Brandt , Prithviraj Dasgupta

Unit-level models for survey data offer many advantages over their area-level counterparts, such as potential for more precise estimates and a natural benchmarking property. However two main challenges occur in this context: accounting for…

统计方法学 · 统计学 2020-05-18 Paul A. Parker , Scott H. Holan , Ryan Janicki

Many data stewards collect confidential data that include fine geography. When sharing these data with others, data stewards strive to disseminate data that are informative for a wide range of spatial and non-spatial analyses while…

统计方法学 · 统计学 2016-02-16 Harrison Quick , Scott H. Holan , Christopher K. Wikle , Jerome P. Reiter

High-resolution estimates of population health indicators are critical for precision public health. We propose a method for high-resolution estimation that fuses distinct data sources: an unbiased, low-resolution data source (e.g.…

统计方法学 · 统计学 2025-08-21 Amy Guan , Marissa Reitsma , Roshni Sahoo , Joshua Salomon , Stefan Wager

We report on our experiences of helping staff of the Scottish Longitudinal Study to create synthetic extracts that can be released to users. In particular, we focus on how the synthesis process can be tailored to produce synthetic extracts…

应用统计 · 统计学 2017-12-13 Gillian M. Raab , Beata Nowok , Chris Dibben

We consider the estimation of densities in multiple subpopulations, where the available sample size in each subpopulation greatly varies. This problem occurs in epidemiology, for example, where different diseases may share similar…

统计方法学 · 统计学 2021-09-15 Jiaming Qiu , Xiongtao Dai , Zhengyuan Zhu

This paper presents a population synthesis model that utilizes the Wasserstein Generative-Adversarial Network (WGAN) for training on incomplete microsamples. By using a mask matrix to represent missing values, the study proposes a WGAN…

机器学习 · 计算机科学 2025-10-02 Tanay Rastogi , Daniel Jonsson , Anders Karlström

The use of synthetic data in machine learning applications and research offers many benefits, including performance improvements through data augmentation, privacy preservation of original samples, and reliable method assessment with fully…

机器学习 · 计算机科学 2026-04-13 Joanna Komorniczak