English
Related papers

Related papers: Identification Risks Evaluation of Partially Synth…

200 papers

Ordinal user-provided ratings across multiple items are frequently encountered in both scientific and commercial applications. Whilst recommender systems are known to do well on these type of data from a predictive point of view, their…

Methodology · Statistics 2025-03-05 Sjoerd Hermes

The large number of publicly available survey datasets of wide variety, albeit useful, raise respondent-level privacy concerns. The synthetic data approach to data privacy and confidentiality has been shown useful in terms of privacy…

Applications · Statistics 2022-05-24 Yixiao Cao , Jingchen Hu

A key challenge for the development and deployment of artificial intelligence (AI) solutions in radiology is solving the associated data limitations. Obtaining sufficient and representative patient datasets with appropriate annotations may…

Image and Video Processing · Electrical Eng. & Systems 2024-07-03 Elena Sizikova , Andreu Badal , Jana G. Delfino , Miguel Lago , Brandon Nelson , Niloufar Saharkhiz , Berkman Sahiner , Ghada Zamzmi , Aldo Badano

We investigate whether generating synthetic data can be a viable strategy for providing access to detailed geocoding information for external researchers, without compromising the confidentiality of the units included in the database. Our…

Applications · Statistics 2020-08-25 Joerg Drechsler , Jingchen Hu

The release of synthetic data generated from a model estimated on the data helps statistical agencies disseminate respondent-level data with high utility and privacy protection. Motivated by the challenge of disseminating sensitive…

Applications · Statistics 2021-02-03 Jingchen Hu , Terrance D. Savitsky

Partial identification approaches are a flexible and robust alternative to standard point-identification approaches in general instrumental variable models. However, this flexibility comes at the cost of a ``curse of cardinality'': the…

Econometrics · Economics 2020-06-30 Florian Gunsilius

Sharing data can often enable compelling applications and analytics. However, more often than not, valuable datasets contain information of a sensitive nature, and thus, sharing them can endanger the privacy of users and organizations. A…

Cryptography and Security · Computer Science 2024-02-28 Emiliano De Cristofaro

False discovery rates (FDR) are an essential component of statistical inference, representing the propensity for an observed result to be mistaken. FDR estimates should accompany observed results to help the user contextualize the relevance…

Methodology · Statistics 2020-10-12 Megan Hollister Murray , Jeffrey D. Blume

Risk is part of the fabric of every business; surprisingly, there is little work on establishing best practices for systematic, repeatable risk identification, arguably the first step of any risk management process. In this paper, we…

General Finance · Quantitative Finance 2015-10-29 Jochen L. Leidner

This study is part of a larger project focused on measuring, understanding, and improving student engagement in programming education. We investigate whether synthetic data generation can help identify at-risk students earlier in a small,…

Computers and Society · Computer Science 2025-05-26 Daniel Flood , Matthew England , Beate Grawemeyer

The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative to real-world data…

Computation and Language · Computer Science 2025-07-25 Tevin Atwal , Chan Nam Tieu , Yefeng Yuan , Zhan Shi , Yuhong Liu , Liang Cheng

We give an overview of eight different software packages and functions available in R for semi- or non-parametric estimation of the hazard rate for right-censored survival data. Of particular interest is the accuracy of the estimation of…

Computation · Statistics 2015-09-11 Yolanda Hagar , Vanja Dukic

We introduce a new class of range restricted formal data privacy standards that condition on owner beliefs about sensitive data ranges. By incorporating this additional information, we can provide a stronger privacy guarantee (e.g. an…

Methodology · Statistics 2026-02-10 Jingchen Hu , Matthew R. Williams , Terrance D. Savitsky

The use of synthetic data provides an opportunity to accelerate online safety research and development efforts while showing potential for bias mitigation, facilitating data storage and sharing, preserving privacy and reducing exposure to…

Computers and Society · Computer Science 2024-02-08 Pica Johansson , Jonathan Bright , Shyam Krishna , Claudia Fischer , David Leslie

Anonymizing microdata requires balancing the reduction of disclosure risk with the preservation of data utility. Traditional evaluations often rely on single measures or two-dimensional risk-utility (R-U) maps, but real-world assessments…

Applications · Statistics 2026-04-10 Oscar Thees , Roman Müller , Matthias Templ

Randomized controlled trials (RCTs) have become powerful tools for assessing the impact of interventions and policies in many contexts. They are considered the gold standard for causal inference in the biomedical fields and many social…

It is shown how to set up, conduct, and analyze large simulation studies with the new R package simsalapar = simulations simplified and launched parallel. A simulation study typically starts with determining a collection of input variables…

Computation · Statistics 2013-09-18 Marius Hofert , Martin Mächler

Bayesian synthetic likelihood (BSL) is a popular method for estimating the parameter posterior distribution for complex statistical models and stochastic processes that possess a computationally intractable likelihood function. Instead of…

Computation · Statistics 2019-07-26 Ziwen An , Leah F South , Christopher Drovandi

This study leverages synthetic data as a validation set to reduce overfitting and ease the selection of the best model in AI development. While synthetic data have been used for augmenting the training set, we find that synthetic data can…

Computer Vision and Pattern Recognition · Computer Science 2023-10-25 Qixin Hu , Alan Yuille , Zongwei Zhou

This article describes techniques employed in the production of a synthetic dataset of driver telematics emulated from a similar real insurance dataset. The synthetic dataset generated has 100,000 policies that included observations about…

Machine Learning · Statistics 2021-02-02 Banghee So , Jean-Philippe Boucher , Emiliano A. Valdez