English
Related papers

Related papers: Improved Inference for Respondent-Driven Sampling …

200 papers

In analyzing big data for finite population inference, it is critical to adjust for the selection bias in the big data. In this paper, we propose two methods of reducing the selection bias associated with the big data sample. The first…

Methodology · Statistics 2019-01-08 Jae Kwang Kim , Zhonglei Wang

Weak signal identification and inference are very important in the area of penalized model selection, yet they are under-developed and not well-studied. Existing inference procedures for penalized estimators are mainly focused on strong…

Methodology · Statistics 2016-11-16 Peibei Shi , Annie Qu

Suppose we are interested in the mean of an outcome that is subject to nonignorable nonresponse. This paper develops new semiparametric estimation methods with instrumental variables which affect nonresponse, but not the outcome. The…

Methodology · Statistics 2024-08-20 Baoluo Sun , Wang Miao , Deshanee S. Wickramarachchi

Frequently, empirical studies are plagued with missing data. When the data are missing not at random, the parameter of interest is not identifiable in general. Without additional assumptions, we can derive bounds of the parameters of…

Methodology · Statistics 2018-09-12 Zhichao Jiang , Peng Ding

We introduce a new predictive mechanism that operates in the presence of hidden confounding across distributionally diverse data sources while ensuring consistent estimation of causal parameters-despite their recognized suboptimality for…

Statistics Theory · Mathematics 2025-04-01 Carlos García Meixide , David Ríos Insua

Random forests is a state-of-the-art supervised machine learning method which behaves well in high-dimensional settings although some limitations may happen when $p$, the number of predictors, is much larger than the number of observations…

Methodology · Statistics 2019-02-01 Louis Capitaine , Robin Genuer , Rodolphe Thiébaut

Social networks play a key role in studying various individual and social behaviors. To use social networks in a study, their structural properties must be measured. For offline social networks, the conventional procedure is…

Social and Information Networks · Computer Science 2018-12-17 Naghmeh Momeni , Michael G. Rabbat

Nonresponse after probability sampling is a universal challenge in survey sampling, often necessitating adjustments to mitigate sampling and selection bias simultaneously. This study explored the removal of bias and effective utilization of…

Methodology · Statistics 2025-11-13 Kosuke Morikawa , Kenji Beppu , Wataru Aida

The two-phase sampling design is a cost-effective strategy widely used in public health research. Analyzing the Phase II sample often involves creating subsample-specific weights. However, these weights can be highly variable, leading to…

Methodology · Statistics 2026-04-07 Xinru Wang , Anyu Zhu , Lauren Kennedy , Abigail Greenleaf , Qixuan Chen

The topic of this paper is prevalence estimation from the perspective of active information. Prevalence among tested individuals has an upward bias under the assumption that individuals' willingness to be tested for the disease increases…

Methodology · Statistics 2022-06-13 Ola Hössjer , Daniel Andrés Díaz-Pachón , Chen Zhao , J. Sunil Rao

Differential privacy is a leading protection setting, focused by design on individual privacy. Many applications, in medical / pharmaceutical domains or social networks, rather posit privacy at a group level, a setting we call integral…

Machine Learning · Statistics 2019-07-04 Hisham Husain , Zac Cranko , Richard Nock

Our aim is to estimate the largest community (a.k.a., mode) in a population composed of multiple disjoint communities. This estimation is performed in a fixed confidence setting via sequential sampling of individuals with replacement. We…

Statistics Theory · Mathematics 2023-09-25 Meera Pai , Nikhil Karamchandani , Jayakrishnan Nair

The network scale-up method (NSUM) is a cost-effective approach to estimating the size or prevalence of a group of people that is hard to reach through a standard survey. The basic NSUM involves two steps: estimating respondents' degrees by…

Methodology · Statistics 2024-01-19 Jessica P. Kunke , Ian Laga , Xiaoyue Niu , Tyler H. McCormick

A central theme in the field of survey statistics is estimating population-level quantities through data coming from potentially non-representative samples of the population. Multilevel Regression and Poststratification (MRP), a model-based…

Methodology · Statistics 2020-07-17 Yuxiang Gao , Lauren Kennedy , Daniel Simpson , Andrew Gelman

Data describing human interactions often suffer from incomplete sampling of the underlying population. As a consequence, the study of contagion processes using data-driven models can lead to a severe underestimation of the epidemic risk.…

Physics and Society · Physics 2015-11-19 Mathieu Génois , Christian L. Vestergaard , Ciro Cattuto , Alain Barrat

Statistical inference with non-probability survey samples is an emerging topic in survey sampling and official statistics and has gained increased attention from researchers and practitioners in the field. Much of the existing literature,…

Methodology · Statistics 2024-10-07 Yang Liu , Meng Yuan , Pengfei Li , Changbao Wu

Pathogen deep-sequencing is an increasingly routinely used technology in infectious disease surveillance. We present a semi-parametric Bayesian Poisson model to exploit these emerging data for inferring infectious disease transmission flows…

Applications · Statistics 2022-01-06 Xiaoyue Xi , Simon EF Spencer , Matthew Hall , M Kate Grabowski , Joseph Kagaayi , Oliver Ratmann

Sequential importance sampling algorithms have been defined to estimate likelihoods in models of ancestral population processes. However, these algorithms are based on features of the models with constant population size, and become…

Statistics Theory · Mathematics 2016-03-24 Coralie Merle , Raphaël Leblois , François Rousset , Pierre Pudlo

This paper addresses the sample selection model within the context of the gender gap problem, where even random treatment assignment is affected by selection bias. By offering a robust alternative free from distributional or specification…

Econometrics · Economics 2024-10-04 Xiaolin Sun , Xueyan Zhao , D. S. Poskitt

Clinical prediction models must be developed using sufficiently large datasets to minimise overfitting and ensure robust predictive performance. Existing sample size calculations assume complete predictor data for all included participants,…