English
Related papers

Related papers: Inference from Small and Big Data Sets with Error …

200 papers

It is increasingly common to collect pre-post data with pseudonyms or self-constructed identifiers. On survey responses from sensitive populations, identifiers may be made optional to encourage higher response rates. The ability to match…

Methodology · Statistics 2023-12-25 Raymond Pomponio , Bailey K. Fosdick , Julia Wrobel , Ryan A. Peterson

This paper examines the foundational concept of random variables in probability theory and statistical inference, demonstrating that their mathematical definition requires no reference to randomization or hypothetical repeated sampling. We…

Other Statistics · Statistics 2025-02-11 Paul W. Vos

We propose a general semi-supervised inference framework focused on the estimation of the population mean. As usual in semi-supervised settings, there exists an unlabeled sample of covariate vectors and a labeled sample consisting of…

Methodology · Statistics 2018-08-15 Anru Zhang , Lawrence D. Brown , T. Tony Cai

Self-normalized processes are basic to many probabilistic and statistical studies. They arise naturally in the the study of stochastic integrals, martingale inequalities and limit theorems, likelihood-based methods in hypothesis testing and…

Probability · Mathematics 2009-09-29 Victor H. de la Peña , Michael J. Klass , Tze Leung Lai

This paper studies transfer learning for estimating the mean of random functions based on discretely sampled data, where, in addition to observations from the target distribution, auxiliary samples from similar but distinct source…

Statistics Theory · Mathematics 2024-03-29 T. Tony Cai , Dongwoo Kim , Hongming Pu

Prediction-powered inference is a framework for performing valid statistical inference when an experimental dataset is supplemented with predictions from a machine-learning system. The framework yields simple algorithms for computing…

Machine Learning · Statistics 2023-11-10 Anastasios N. Angelopoulos , Stephen Bates , Clara Fannjiang , Michael I. Jordan , Tijana Zrnic

The "large p, small n" paradigm arises in microarray studies, where expression levels of thousands of genes are monitored for a small number of subjects. There has been an increasing demand for study of asymptotics for the various…

Statistics Theory · Mathematics 2007-06-13 Michael R. Kosorok , Shuangge Ma

Random sampling is an essential tool in the processing and transmission of data. It is used to summarize data too large to store or manipulate and meet resource constraints on bandwidth or battery power. Estimators that are applied to the…

Databases · Computer Science 2015-03-19 Edith Cohen , Haim Kaplan

A large class of problems in sciences and engineering can be formulated as the general problem of constructing random intervals with pre-specified coverage probabilities for the mean. Wee propose a general approach for statistical inference…

Statistics Theory · Mathematics 2013-06-11 Xinjia Chen

Consider the task of generating samples from a tilted distribution of a random vector whose underlying distribution is unknown, but samples from it are available. This finds applications in fields such as finance and climate science, and in…

In this paper, a new randomized response technique aimed at protecting respondents' privacy is proposed. It is designed for estimating the population total, or the population mean, of a quantitative characteristic. It provides a~high degree…

Methodology · Statistics 2021-01-25 Jaromir Antoch , Francesco Mola , Ondrej Vozar

Fitting linear regression models can be computationally very expensive in large-scale data analysis tasks if the sample size and the number of variables are very large. Random projections are extensively used as a dimension reduction tool…

Statistics Theory · Mathematics 2017-01-20 Gian-Andrea Thanei , Christina Heinze , Nicolai Meinshausen

Distorted distributions were introduced in the context of actuarial science for several variety of insurance problems. In this paper we consider the quantile-based probabilistic mean value theorem given in Di Crescenzo et al. [4] and…

Probability · Mathematics 2025-01-03 Antonio Di Crescenzo , Barbara Martinucci , Julio Mulero

The exploration of associations between random objects with complex geometric structures has catalyzed the development of various novel statistical tests encompassing distance-based and kernel-based statistics. These methods have various…

Methodology · Statistics 2024-10-28 Zhe Gao , Roulin Wang , Xueqin Wang , Heping Zhang

Introductory statistical inference texts and courses treat the point estimation, hypothesis testing, and interval estimation problems separately, with primary emphasis on large-sample approximations. Here I present an alternative approach…

Other Statistics · Statistics 2017-07-14 Ryan Martin

The increased availability of massive data sets provides a unique opportunity to discover subtle patterns in their distributions, but also imposes overwhelming computational challenges. To fully utilize the information contained in big…

Statistics Theory · Mathematics 2018-04-12 Stanislav Volgushev , Shih-Kang Chao , Guang Cheng

The term moderate deviations is often used in the literature to mean a class of large deviation principles that, in some sense, fills the gap between a convergence in probability of some random variables to a constant and a weak convergence…

Probability · Mathematics 2024-11-20 Rita Giuliano , Claudio Macci , Barbara Pacchiarotti

We study linear regressions in a context where the outcome of interest and some of the covariates are observed in two different datasets that cannot be matched. Traditional approaches obtain point identification by relying, often…

Econometrics · Economics 2025-11-18 Xavier D'Haultfoeuille , Christophe Gaillac , Arnaud Maurel

Randomization inference is a widely-used and appealing approach for analyzing treatment effects in randomized experiments, as it is finite-sample valid and does not require any distributional assumptions. However, naive application of…

Econometrics · Economics 2026-05-12 Xinran Li , Peizan Sheng , Zeyang Yu

We study inference with a small labeled sample, a large unlabeled sample, and high-quality predictions from an external model. We link prediction-powered inference with empirical likelihood by stacking supervised estimating equations based…

Methodology · Statistics 2025-12-19 Guanghui Wang , Mengtao Wen , Changliang Zou