English
Related papers

Related papers: Toward design-based inference for data integration

200 papers

Sparsity in a regression context makes the model itself an object of interest, pointing to a confidence set of models as the appropriate presentation of evidence. A difficulty in areas such as genomics, where the number of candidate…

Statistics Theory · Mathematics 2026-02-24 Heather Battey , Daniel Garcia Rasines , Yanbo Tang

In scientific applications, multivariate observations often come in tandem with temporal or spatial covariates, with which the underlying signals vary smoothly. The standard approaches such as principal component analysis and factor…

Statistics Theory · Mathematics 2019-10-15 Mark Koudstaal , Dengdeng Yu , Dehan Kong , Fang Yao

In the analysis of observational data in social sciences and businesses, it is difficult to obtain a "(quasi) single-source dataset" in which the variables of interest are simultaneously observed. Instead, multiple-source datasets are…

Methodology · Statistics 2021-09-02 Masaki Mitsuhiro , Takahiro Hoshino

We establish a mathematical framework that formally validates the two-phase ``super-population viewpoint'' proposed by Hartley and Sielken [Biometrics 31 (1975) 411--422] by defining a product probability space which includes both the…

Statistics Theory · Mathematics 2007-06-13 Susana Rubin-Bleuer , Ioana Schiopu Kratina

Nonprobability (convenience) samples are increasingly sought to stabilize estimations for one or more population variables of interest that are performed using a randomized survey (reference) sample by increasing the effective sample size.…

We study regression discontinuity designs when covariates are included in the estimation. We examine local polynomial estimators that include discrete or continuous covariates in an additive separable way, but without imposing any…

Econometrics · Economics 2019-07-02 Sebastian Calonico , Matias D. Cattaneo , Max H. Farrell , Rocio Titiunik

The following paper presents nonprobsvy -- an R package for inference based on non-probability samples. The package implements various approaches that can be categorized into three groups: prediction-based approach, inverse probability…

Methodology · Statistics 2025-08-21 Łukasz Chrostowski , Piotr Chlebicki , Maciej Beręsewicz

In experimental design, we are given a large collection of vectors, each with a hidden response value that we assume derives from an underlying linear model, and we wish to pick a small subset of the vectors such that querying the…

Machine Learning · Computer Science 2019-02-05 Michał Dereziński , Kenneth L. Clarkson , Michael W. Mahoney , Manfred K. Warmuth

In modern experimental science, there is a common problem of estimating the coefficients of a linear regression in a context where the variables of interest cannot be observed simultaneously. When there is a categorical variable that is…

Methodology · Statistics 2025-03-10 Polina Arsenteva , Mohamed Amine Benadjaoud , Hervé Cardot

Recommender systems often suffer from selection bias as users tend to rate their preferred items. The datasets collected under such conditions exhibit entries missing not at random and thus are not randomized-controlled trials representing…

Information Retrieval · Computer Science 2024-03-05 Wonbin Kweon , Hwanjo Yu

We examine study designs for extending (generalizing or transporting) causal inferences from a randomized trial to a target population. Specifically, we consider nested trial designs, where randomized individuals are nested within a sample…

In experimental causal inference, we distinguish between two sources of uncertainty: design uncertainty, due to the treatment assignment mechanism, and sampling uncertainty, when the sample is drawn from a super-population. This distinction…

This paper studies inference in randomized controlled trials with covariate-adaptive randomization when there are multiple treatments. More specifically, we study inference about the average effect of one or more treatments relative to…

Econometrics · Economics 2019-01-21 Federico A. Bugni , Ivan A. Canay , Azeem M. Shaikh

Commonly used methods to analyze incomplete longitudinal clinical trial data include complete case analysis (CC) and last observation carried forward (LOCF). However, such methods rest on strong assumptions, including missing completely at…

Statistics Theory · Mathematics 2007-06-13 Ivy Jansen , Caroline Beunckens , Geert Molenberghs , Geert Verbeke , Craig Mallinckrodt

In various applications of regression analysis, in addition to errors in the dependent observations also errors in the predictor variables play a substantial role and need to be incorporated in the statistical modeling process. In this…

Statistics Theory · Mathematics 2020-09-03 Katharina Proksch , Nicolai Bissantz , Hajo Holzmann

Surveys usually suffer from non-response, which decreases the effective sample size. Item non-response is typically handled by means of some form of random imputation if we wish to preserve the distribution of the imputed variable. This…

Methodology · Statistics 2017-08-04 Guillaume Chauvet , Wilfried Do Paco

This paper studies inference in randomized controlled trials with multiple treatments, where treatment status is determined according to a "matched tuples" design. Here, by a matched tuples design, we mean an experimental design where units…

Econometrics · Economics 2023-11-06 Yuehao Bai , Jizhou Liu , Max Tabord-Meehan

Seemingly unrelated linear regression models are introduced in which the distribution of the errors is a finite mixture of Gaussian components. Identifiability conditions are provided. The score vector and the Hessian matrix are derived.…

Methodology · Statistics 2014-03-18 Giuliano Galimberti , Elena Scardovi , Gabriele Soffritti

This article introduces a novel nonparametric methodology for Generalized Linear Models which combines the strengths of the binary regression and latent variable formulations for categorical data, while overcoming their disadvantages.…

Machine Learning · Statistics 2021-10-12 K. P. Chowdhury

Missing data is frequently encountered in many areas of statistics. Propensity score weighting is a popular method for handling missing data. The propensity score method employs a response propensity model, but correct specification of the…

Methodology · Statistics 2024-03-28 Hengfang Wang , Jae Kwang Kim , Jeongseop Han , Youngjo Lee