中文
相关论文

相关论文: Fully Synthetic Data for Complex Surveys

200 篇论文

Multi-level modeling is an important approach for analyzing complex survey data using multi-stage sampling. However, estimation of multi-level models can be challenging when we combine several datasets with distinct hierarchies with…

统计方法学 · 统计学 2023-09-26 Seho Park , A James OMalley

Generating synthetic data, with or without differential privacy, has attracted significant attention as a potential solution to the dilemma between making data easily available, and the privacy of data subjects. Several works have shown…

统计方法学 · 统计学 2023-11-01 Ossi Räisä , Joonas Jälkö , Antti Honkela

The public availability of collections containing user preferences is of vital importance for performing offline evaluations in the field of recommender systems. However, the number of rating datasets is limited because of the costs…

信息检索 · 计算机科学 2019-09-04 Diego Monti , Giuseppe Rizzo , Maurizio Morisio

Household and individual-level sociodemographic data are essential for understanding human-infrastructure interaction and policymaking. However, the Public Use Microdata Sample (PUMS) offers only a sample at the state level, while census…

机器学习 · 计算机科学 2024-07-03 Xiao Qian , Utkarsh Gangwal , Shangjia Dong , Rachel Davidson

We introduce a constraint-programming framework for generating synthetic populations that reproduce target statistics with high precision while enforcing full individual consistency. Unlike data-driven approaches that infer distributions…

机器学习 · 统计学 2025-12-09 Thierry Petit , Arnault Pachot

The idea to generate synthetic data as a tool for broadening access to sensitive microdata has been proposed for the first time three decades ago. While first applications of the idea emerged around the turn of the century, the approach…

密码学与安全 · 计算机科学 2023-04-06 Joerg Drechsler , Anna-Carolina Haensch

Differential privacy allows quantifying privacy loss resulting from accessing sensitive personal data. Repeated accesses to underlying data incur increasing loss. Releasing data as privacy-preserving synthetic data would avoid this…

机器学习 · 统计学 2021-06-10 Joonas Jälkö , Eemil Lagerspetz , Jari Haukka , Sasu Tarkoma , Antti Honkela , Samuel Kaski

Introduction: The amount of data generated by original research is growing exponentially. Publicly releasing them is recommended to comply with the Open Science principles. However, data collected from human participants cannot be released…

机器学习 · 统计学 2023-10-11 Rémy Chapelle , Bruno Falissard

Access to individual-level health data is essential for gaining new insights and advancing science. In particular, modern methods based on artificial intelligence rely on the availability of and access to large datasets. In the health…

Synthetic population generation is the process of combining multiple socioeconomic and demographic datasets from different sources and/or granularity levels, and downscaling them to an individual level. Although it is a fundamental step for…

机器学习 · 计算机科学 2019-11-12 Colin Wan , Zheng Li , Alicia Guo , Yue Zhao

We present new techniques for automatically constructing probabilistic programs for data analysis, interpretation, and prediction. These techniques work with probabilistic domain-specific data modeling languages that capture key properties…

Bayesian analysis is increasingly popular for use in social science and other application areas where the data are observations from an informative sample. An informative sampling design leads to inclusion probabilities that are correlated…

统计理论 · 数学 2016-06-07 Terrance D. Savitsky , Daniell Toth

We describe results on the creation and use of synthetic data that were derived in the context of a project to make synthetic extracts available for users of the UK Longitudinal Studies. A critical review of existing methods of inference…

统计方法学 · 统计学 2017-12-12 Gillian Raab , Beata Nowok , Chris Dibben

We propose a categorical data synthesizer with a quantifiable disclosure risk. Our algorithm, named Perturbed Gibbs Sampler, can handle high-dimensional categorical data that are often intractable to represent as contingency tables. The…

机器学习 · 统计学 2013-12-20 Yubin Park , Joydeep Ghosh

In recent years hypergraphs have emerged as a powerful tool to study systems with multi-body interactions which cannot be trivially reduced to pairs. While highly structured methods to generate synthetic data have proved fundamental for the…

社会与信息网络 · 计算机科学 2024-10-10 Nicolò Ruggeri , Federico Battiston , Caterina De Bacco

The integration of data from multiple sources is increasingly used to achieve larger sample sizes and enhance population diversity. Our previous work established that, under random sampling from the same underlying population, integrating…

统计方法学 · 统计学 2026-01-01 Farimah Shamsi , Andriy Derkach

Multiple data sources are becoming increasingly available for statistical analyses in the era of big data. As an important example in finite-population inference, we consider an imputation approach to combining a probability sample with big…

统计方法学 · 统计学 2018-07-10 Shu Yang , Jae Kwang Kim

As more tech companies engage in rigorous economic analyses, we are confronted with a data problem: in-house papers cannot be replicated due to use of sensitive, proprietary, or private data. Readers are left to assume that the obscured…

综合经济学 · 经济学 2020-11-10 Allison Koenecke , Hal Varian

Non-representative surveys are commonly used and widely available but suffer from selection bias that generally cannot be entirely eliminated using weighting techniques. Instead, we propose a Bayesian method to synthesize longitudinal…

统计方法学 · 统计学 2024-07-08 Nathaniel Dyrkton , Paul Gustafson , Harlan Campbell

This paper introduces smoothed pseudo-population bootstrap methods for the purposes of variance estimation and the construction of confidence intervals for finite population quantiles. In an i.i.d. context, it has been shown that resampling…

统计方法学 · 统计学 2025-09-30 Vanessa McNealis , Christian Léger