中文
相关论文

相关论文: On application of a response propensity model to e…

200 篇论文

For massive data, the family of subsampling algorithms is popular to downsize the data volume and reduce computational burden. Existing studies focus on approximating the ordinary least squares estimate in linear regression, where…

统计计算 · 统计学 2019-06-27 HaiYing Wang , Rong Zhu , Ping Ma

Selection bias is a serious potential problem for inference about relationships of scientific interest based on samples without well-defined probability sampling mechanisms. Motivated by the potential for selection bias in (a) estimated…

Logistic regression is the most commonly used method for constructing predictive models for binary responses. One significant drawback to this approach, however, is that the asymptotes of the logistic response function are fixed at 0 and 1,…

统计方法学 · 统计学 2026-02-09 Anthony Almudevar , Jacob Almudevar

In many surveys inexpensive auxiliary variables are available that can help us to make more precise estimation about the main variable. Using auxiliary variable has been extended by regression estimators for rare and cluster populations. In…

统计理论 · 数学 2018-03-14 Bardia Panahbehagh , Afshin Parvardeh , Babak Mohammadi

Marginal imputation, which consists of imputing each item requiring imputation separately, is often used in surveys. This type of imputation procedures leads to asymptotically unbiased estimators of simple parameters such as population…

统计方法学 · 统计学 2015-11-04 Hélène Chaput , Guillaume Chauvet , David Haziza , Laurianne Salembier , Julie Solard

Population size estimates for hidden and hard-to-reach populations are particularly important when members are known to suffer from disproportion health issues or to pose health risks to the larger ambient population in which they are…

社会与信息网络 · 计算机科学 2018-07-04 Bilal Khan , Hsuan-Wei Lee , Ian Fellows , Kirk Dombrowski

Imputing missing values is an important preprocessing step in data analysis, but the literature offers little guidance on how to choose between different imputation models. This letter suggests adopting the imputation model that generates a…

统计方法学 · 统计学 2021-07-13 Moritz Marbach

This work is concerned with the estimation of hard-to-reach population sizes using a single respondent-driven sampling (RDS) survey, a variant of chain-referral sampling that leverages social relationships to reach members of a hidden…

In analyzing big data for finite population inference, it is critical to adjust for the selection bias in the big data. In this paper, we propose two methods of reducing the selection bias associated with the big data sample. The first…

统计方法学 · 统计学 2019-01-08 Jae Kwang Kim , Zhonglei Wang

Propensity score weighting is a common method for estimating treatment effects with survey data. The method is applied to minimize confounding using measured covariates that are often different between individuals in treatment and control.…

统计方法学 · 统计学 2026-02-06 Yukang Zeng , Fan Li , Guangyu Tong

Respondent-driven sampling (RDS) is an approach to sampling design and analysis which utilizes the networks of social relationships that connect members of the target population, using chain-referral methods to facilitate sampling. RDS…

统计方法学 · 统计学 2015-08-19 Yakir Berchenko , Jonathan Rosenblatt , Simon D. W. Frost

We investigate the complexity of logistic regression models which is defined by counting the number of indistinguishable distributions that the model can represent (Balasubramanian, 1997). We find that the complexity of logistic models with…

机器学习 · 统计学 2019-03-04 Nicola Bulso , Matteo Marsili , Yasser Roudi

Statistical causal inference from observational studies often requires adjustment for a possibly multi-dimensional variable, where dimension reduction is crucial. The propensity score, first introduced by Rosenbaum and Rubin, is a popular…

统计理论 · 数学 2020-04-28 Hui Guo , Philip Dawid , Giovanni Berzuini

A solution to control for nonresponse bias consists of multiplying the design weights of respondents by the inverse of estimated response probabilities to compensate for the nonrespondents. Maximum likelihood and calibration are two…

统计方法学 · 统计学 2023-10-27 Caren Hasler

The usage of machine learning methods in traditional surveys including official statistics, is still very limited. Therefore, we propose a predictor supported by these algorithms, which can be used to predict any population or subpopulation…

统计方法学 · 统计学 2025-07-14 Tomasz Żądło , Adam Chwila

We consider the fundamental problem of how to automatically construct summary statistics for implicit generative models where the evaluation of the likelihood function is intractable, but sampling data from the model is possible. The idea…

机器学习 · 统计学 2021-03-31 Yanzhi Chen , Dinghuai Zhang , Michael Gutmann , Aaron Courville , Zhanxing Zhu

When random effects are correlated with sample design variables, the usual approach of employing individual survey weights (constructed to be inversely proportional to the unit survey inclusion probabilities) to form a pseudo-likelihood no…

统计方法学 · 统计学 2021-08-26 Terrance D. Savitsky , Matthew R. Williams

F\'elix-Medina and Thompson (2004) proposed a variant of link-tracing sampling to estimate the size of a hidden population such as drug users, sexual workers or homeless people. In their variant a sampling frame of sites where the members…

统计方法学 · 统计学 2015-06-23 Martin H. Félix Medina

Respondent-driven sampling is a widely-used network sampling technique, designed to sample from hard-to-reach populations. Estimation from the resulting samples is an area of active research, with software available to compute at least four…

应用统计 · 统计学 2010-12-21 Amber Tomas , Krista J. Gile

There is a growing trend among statistical agencies to explore non-probability data sources for producing more timely and detailed statistics, while reducing costs and respondent burden. Coverage and measurement error are two issues that…

统计方法学 · 统计学 2024-09-19 Lyndon Ang , Robert Clark , Bronwyn Loong , Anders Holmberg