English
Related papers

Related papers: Thresholding Nonprobability Units in Combined Data…

200 papers

Integrating probability and non-probability samples is increasingly important, yet unknown sampling mechanisms in non-probability sources complicate identification and efficient estimation. We develop semiparametric theory for dual-frame…

Methodology · Statistics 2026-01-14 Kosuke Morikawa , Jae Kwang Kim

In this work, we revisit the problem of active sequential prediction-powered mean estimation, where at each round one must decide the query probability of the ground-truth label upon observing the covariates of a sample. Furthermore, if the…

Machine Learning · Statistics 2026-04-21 Maria-Eleni Sfyraki , Jun-Kun Wang

Often, government agencies and survey organizations know the population counts or percentages for some of the variables in a survey. These may be available from auxiliary sources, for example, administrative databases or other high quality…

Methodology · Statistics 2020-11-12 Olanrewaju Akande , Gabriel Madson , D. Sunshine Hillygus , Jerome P. Reiter

Multivariate meta-analysis of test accuracy studies when tests are evaluated in terms of sensitivity and specificity at more than one threshold represents an effective way to synthesize results by fully exploiting the data, if compared to…

Methodology · Statistics 2019-01-29 Annamaria Guolo , Duc Khanh To

In many applications, data cluster. Failing to take the cluster structure into consideration generally leads to underestimated variances of point estimators and inflated type I errors in hypothesis tests. Many circumstance-dependent…

Methodology · Statistics 2025-07-21 Jiahua Chen , Pengfei Li , Yukun Liu , James V. Zidek

We devise survey-weighted pseudo posterior distribution estimators under two-stage informative sampling of both primary clusters and secondary nested units for a one-way analysis of variance (ANOVA) population generating model as a simple…

Methodology · Statistics 2023-05-16 Terrance D. Savitsky , Matthew R. Williams , Sanvesh Srivastava

Propensity score weighting is a tool for causal inference to adjust for measured confounders. Survey data are often collected under complex sampling designs such as multistage cluster sampling, which presents challenges for propensity score…

Methodology · Statistics 2016-07-27 Shu Yang

Statistical estimation and inference for marginal hazard models with varying coefficients for multivariate failure time data are important subjects in survival analysis. A local pseudo-partial likelihood procedure is proposed for estimating…

Statistics Theory · Mathematics 2009-09-29 Jianwen Cai , Jianqing Fan , Haibo Zhou , Yong Zhou

In this paper we present a frequentist-Bayesian hybrid method for estimating covariances of unfolded distributions using pseudo-experiments. The method is compared with other covariance estimation methods using the unbiased Rao-Cramer bound…

Methodology · Statistics 2021-10-19 Pim Jordi Verschuuren

We consider parameter estimation of stochastic differential equations driven by a Wiener process and a compound Poisson process as small noises. The goal is to give a threshold-type quasi-likelihood estimator and show its consistency and…

Statistics Theory · Mathematics 2023-12-20 Mitsuki Kobayashi , Yasutaka Shimizu

Least squares estimators, when trained on a few target domain samples, may predict poorly. Supervised domain adaptation aims to improve the predictive accuracy by exploiting additional labeled training samples from a source distribution…

Machine Learning · Computer Science 2021-06-02 Bahar Taskesen , Man-Chung Yue , Jose Blanchet , Daniel Kuhn , Viet Anh Nguyen

We study the sample complexity of estimating the covariance matrix $T$ of a distribution $\mathcal{D}$ over $d$-dimensional vectors, under the assumption that $T$ is Toeplitz. This assumption arises in many signal processing problems, where…

Signal Processing · Electrical Eng. & Systems 2019-10-31 Yonina C. Eldar , Jerry Li , Cameron Musco , Christopher Musco

We consider the problem of parametric statistical inference when likelihood computations are prohibitively expensive but sampling from the model is possible. Several so-called likelihood-free methods have been developed to perform inference…

Machine Learning · Statistics 2020-09-14 Owen Thomas , Ritabrata Dutta , Jukka Corander , Samuel Kaski , Michael U. Gutmann

Decision theory does not traditionally include uncertainty over utility functions. We argue that the a person's utility value for a given outcome can be treated as we treat other domain attributes: as a random variable with a density…

Artificial Intelligence · Computer Science 2013-01-18 Urszula Chajewska , Daphne Koller

We study the problem of reconstructing the probability measure of the Curie-Weiss model from a sample of the voting behaviour of a subset of the population. While originally used to study phase transitions in statistical mechanics, the…

Probability · Mathematics 2025-08-06 Miguel Ballesteros , Ivan Naumkin , Gabor Toth

Two-phase sampling designs are frequently employed in epidemiological studies and large-scale health surveys. In such designs, certain variables are exclusively collected within a second-phase random subsample of the initial first-phase…

Methodology · Statistics 2024-03-25 Lingxiao Wang

Given additional distributional information in the form of moment restrictions, kernel density and distribution function estimators with implied generalised empirical likelihood probabilities as weights achieve a reduction in variance due…

Methodology · Statistics 2019-10-08 Vitaliy Oryshchenko , Richard J. Smith

Model-assisted estimation with complex survey data is an important practical problem in survey sampling. When there are many auxiliary variables, selecting significant variables associated with the study variable would be necessary to…

Methodology · Statistics 2020-04-01 Shonosuke Sugasawa , Jae Kwang Kim

Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is fundamental to science, public health, and online trust and safety. Standard approaches…

Artificial Intelligence · Computer Science 2026-04-24 Fridolin Linder , Thomas Leeper , Daniel Haimovich , Niek Tax , Lorenzo Perini , Milan Vojnovic

Preferential sampling has attracted considerable attention in geostatistics since the pioneering work of Diggle et al. (2010). A variety of likelihood-based approaches have been developed to correct estimation bias by explicitly modelling…

Methodology · Statistics 2025-11-06 Changqing Lu , Ganggang Xu , Junho Yang , Yongtao Guan