English
Related papers

Related papers: Bayesian Data Synthesis and the Utility-Risk Trade…

200 papers

To combat the HIV/AIDS pandemic effectively, targeted interventions among certain key populations play a critical role. Examples of such key populations include sex workers, people who inject drugs, and men who have sex with men. While…

Applications · Statistics 2021-07-14 Jacob Parsons , Xiaoyue Niu , Le Bao

We propose a categorical data synthesizer with a quantifiable disclosure risk. Our algorithm, named Perturbed Gibbs Sampler, can handle high-dimensional categorical data that are often intractable to represent as contingency tables. The…

Machine Learning · Statistics 2013-12-20 Yubin Park , Joydeep Ghosh

One of the major research questions regarding human microbiome studies is the feasibility of designing interventions that modulate the composition of the microbiome to promote health and cure disease. This requires extensive understanding…

Methodology · Statistics 2021-11-18 Matthew D. Koslovsky , Kristi L. Hoffman , Carrie R. Daniel , Marina Vannucci

Many datasets describing contacts in a population suffer from incompleteness due to population sampling and underreporting of contacts. Data-driven simulations of spreading processes using such incomplete data lead to an underestimation of…

Physics and Society · Physics 2017-09-07 Julie Fournet , Alain Barrat

We introduce a Bayesian approach for analyzing (possibly) high-dimensional dependent data that are distributed according to a member from the natural exponential family of distributions. This problem requires extensive methodological…

Methodology · Statistics 2019-04-19 Jonathan R. Bradley , Scott H. Holan , Christopher K. Wikle

The purpose of this study is to leverage modern technology (such as mobile or web apps in Beckman et al. (2014)) to enrich epidemiology data and infer the transmission of disease. Homogeneity related research on population level has been…

Applications · Statistics 2015-09-02 Kai Fan , Allison E. Aiello , Katherine A. Heller

Learning the structure of Bayesian networks from data provides insights into underlying processes and the causal relationships that generate the data, but its usefulness depends on the homogeneity of the data population, a condition often…

To develop public health intervention models using microsimulations, extensive personal information about inhabitants is needed, such as socio-demographic, economic and health figures. Data confidentiality is an essential characteristic of…

Applications · Statistics 2022-02-10 M. A. Nicolaie , Koen Fussenich , Caroline Ameling , Hendriek C. Boshuizen

Finite mixture model is an important branch of clustering methods and can be applied on data sets with mixed types of variables. However, challenges exist in its applications. First, it typically relies on the EM algorithm which could be…

Machine Learning · Statistics 2019-05-10 Shu Wang , Jonathan G. Yabes , Chung-Chou H. Chang

In the current data driven era, synthetic data, artificially generated data that resembles the characteristics of real world data without containing actual personal information, is gaining prominence. This is due to its potential to…

Machine Learning · Computer Science 2023-09-06 Tshilidzi Marwala , Eleonore Fournier-Tombs , Serge Stinckwich

An informative sampling design leads to the selection of units whose inclusion probabilities are correlated with the response variable of interest. Model inference performed on the resulting observed sample will be biased for the population…

Methodology · Statistics 2018-06-29 Matthew R. Williams , Terrance D. Savitsky

Enhancing reproducibility and data accessibility is essential to scientific research. However, ensuring data privacy while achieving these goals is challenging, especially in the medical field, where sensitive data are often commonplace.…

Methodology · Statistics 2025-09-24 Marta Cipriani , Lorenzo Di Rocco , Maria Puopolo , Marco Alfò

Missing data is a common concern in health datasets, and its impact on good decision-making processes is well documented. Our study's contribution is a methodology for tackling missing data problems using a combination of synthetic dataset…

Machine Learning · Computer Science 2022-11-08 Gift Khangamwa , Terence L. van Zyl , Clint J. van Alten

The analysis of data from multiple experiments, such as observations of several individuals, is commonly approached using mixed-effects models, which account for variation between individuals through hierarchical representations. This makes…

Computation · Statistics 2026-03-05 Henrik Häggström , Sebastian Persson , Marija Cvijovic , Umberto Picchini

We introduce the SoftBart approach from Bayesian ensemble learning to estimate the relationship between multipollutant mixtures and health on chronic exposures in epidemiology research. This approach offers several key advantages over…

Quantitative Methods · Quantitative Biology 2025-05-26 Yu-Chien Ning , Xin Zhou , Francine Laden , Molin Wang

Ideally, a meta-analysis will summarize data from several unbiased studies. Here we consider the less than ideal situation in which contributing studies may be compromised by measurement error. Measurement error affects every study design,…

This paper evaluates synthetically generated healthcare data for biases and investigates the effect of fairness mitigation techniques on utility-fairness. Privacy laws limit access to health data such as Electronic Medical Records (EMRs) to…

Machine Learning · Computer Science 2022-03-10 Karan Bhanot , Ioana Baldini , Dennis Wei , Jiaming Zeng , Kristin P. Bennett

During an epidemic, the information available to individuals in the society deeply influences their belief of the epidemic spread, and consequently the preventive measures they take to stay safe from the infection. In this paper, we develop…

Systems and Control · Electrical Eng. & Systems 2022-07-26 Shraddha Pathak , Ankur A. Kulkarni

Synthetic datasets are often presented as a silver-bullet solution to the problem of privacy-preserving data publishing. However, for many applications, synthetic data has been shown to have limited utility when used to train predictive…

The acute phase of the Covid-19 pandemic has made apparent the need for decision support based upon accurate epidemic modeling. This process is substantially hampered by under-reporting of cases and related data incompleteness issues. In…

Applications · Statistics 2026-03-10 Anastasios Apsemidis , Nikolaos Demiris
‹ Prev 1 3 4 5 6 7 10 Next ›