English
Related papers

Related papers: Efficiency Gains from Using Auxiliary Variables in…

200 papers

We consider a general statistical estimation problem involving a finite-dimensional target parameter vector. Beyond an internal data set drawn from the population distribution, external information, such as additional individual data or…

Methodology · Statistics 2025-07-31 Guorong Dai , Lingxuan Shao , Jinbo Chen

Generalised regression estimation allows one to make use of available auxiliary information in survey sampling. We develop three types of generalised regression estimator when the auxiliary data cannot be matched perfectly to the sample…

Methodology · Statistics 2020-05-20 Li-Chun Zhang

The era of huge data necessitates highly efficient machine learning algorithms. Many common machine learning algorithms, however, rely on computationally intensive subroutines that are prohibitively expensive on large datasets. Oftentimes,…

Machine Learning · Computer Science 2023-09-26 Mo Tiwari

The odds ratio measure is used in health and social surveys where the odds of a certain event is to be compared between two populations. It is defined using logistic regression, and requires that data from surveys are accompanied by their…

Methodology · Statistics 2014-07-01 C. Goga , A Ruiz-Gazen

It is widely admitted that structured nonparametric modeling that circumvents the curse of dimensionality is important in nonparametric estimation. In this paper we show that the same holds for semi-parametric estimation. We argue that…

Statistics Theory · Mathematics 2011-04-25 Kyusang Yu , Enno Mammen , Byeong U. Park

We aim to make inferences about a smooth, finite-dimensional parameter by fusing data from multiple sources together. Previous works have studied the estimation of a variety of parameters in similar data fusion settings, including in the…

Methodology · Statistics 2025-02-03 Sijia Li , Alex Luedtke

Model-assisted estimation with complex survey data is an important practical problem in survey sampling. When there are many auxiliary variables, selecting significant variables associated with the study variable would be necessary to…

Methodology · Statistics 2020-04-01 Shonosuke Sugasawa , Jae Kwang Kim

Optimization problems with an auxiliary latent variable structure in addition to the main model parameters occur frequently in computer vision and machine learning. The additional latent variables make the underlying optimization task…

Machine Learning · Computer Science 2020-03-13 Christopher Zach , Huu Le

Estimation and inference in dynamic discrete choice models often relies on approximation to lower the computational burden of dynamic programming. Unfortunately, the use of approximation can impart substantial bias in estimation and results…

Econometrics · Economics 2020-10-23 Ben Deaner

We present a new method in problems where estimates are needed for finite population domains with small or even zero sample sizes. In contrast to known estimation methods, an auxiliary information is used to model sizes of population units…

Statistics Theory · Mathematics 2014-06-23 Andrius Čiginas , Tomas Rudys

Marginal model is a popular instrument for studying longitudinal data and cluster data. This paper investigates the estimator of marginal model with subgroup auxiliary information. To marginal model, we propose a new type of auxiliary…

Methodology · Statistics 2018-06-11 Jie He , Xiaogang Duan , Shumei Zhang , Hui Li

Estimating causal effects from randomized experiments is central to clinical research. Reducing the statistical uncertainty in these analyses is an important objective for statisticians. Registries, prior trials, and health records…

Machine Learning · Statistics 2021-12-06 Alejandro Schuler , David Walsh , Diana Hall , Jon Walsh , Charles Fisher

Statistical inference is considered for variables of interest, called primary variables, when auxiliary variables are observed along with the primary variables. We consider the setting of incomplete data analysis, where some primary…

Methodology · Statistics 2019-03-27 Shinpei Imori , Hidetoshi Shimodaira

Annotated datasets are an essential ingredient to train, evaluate, compare and productionalize supervised machine learning models. It is therefore imperative that annotations are of high quality. For their creation, good quality management…

Machine Learning · Computer Science 2024-05-30 Jan-Christoph Klie , Juan Haladjian , Marc Kirchner , Rahul Nair

Numerical solutions to fractional differential equations can be extremely computationally intensive due to the effect of non-local derivatives in which all previous time points contribute to the current iteration. In finite difference…

Mathematical Physics · Physics 2010-04-30 Brian P. Sprouse , Christopher L. MacDonald , Gabriel A. Silva

This paper investigates the role of the augmentation parameter in the Finite Selection Model (FSM) and its impact on estimator performance. Through a comprehensive Monte Carlo simulation study, we analyze the sensitivity of bias, variance,…

Methodology · Statistics 2026-03-09 Safaa K. Kadhem

Ensembling methods are well known for improving prediction accuracy. However, they are limited in the sense that they cannot discriminate among component models effectively. In this paper, we propose stacking with auxiliary features that…

Computation and Language · Computer Science 2016-05-30 Nazneen Fatema Rajani , Raymond J. Mooney

This paper proposes a class of ratio type estimators of finite population variance, when the population variance of an auxiliary character is known. Asymptotic expression for mean square error (MSE) is derived and compared with the mean…

Statistics Theory · Mathematics 2013-11-27 Jayant Singh , Viplav K. Singh , Sachin Malik , Rajesh Singh

There has been increasing interest in recent years in the development of approaches to estimate causal effects when the number of potential confounders is prohibitively large. This growth in interest has led to a number of potential…

Methodology · Statistics 2020-02-05 Joseph Antonelli , Matthew Cefalu

Theoretical guarantees for causal inference using propensity scores are partly based on the scores behaving like conditional probabilities. However, scores between zero and one, especially when outputted by flexible statistical estimators,…

Methodology · Statistics 2024-11-12 Rom Gutman , Ehud Karavani , Yishai Shimoni