English
Related papers

Related papers: A note on stratification errors in the analysis of…

200 papers

The past two decades have witnessed a surge of new research in the analysis of randomized experiments. The emergence of this literature may seem surprising given the widespread use and long history of experiments as the "gold standard" in…

Econometrics · Economics 2025-04-03 Yuehao Bai , Azeem M. Shaikh , Max Tabord-Meehan

Selection bias arises when the probability that an observation enters a dataset depends on variables related to the quantities of interest, leading to systematic distortions in estimation and uncertainty quantification. For example, in…

Machine learning algorithms are increasingly used to inform critical decisions. There is a growing concern about bias, that algorithms may produce uneven outcomes for individuals in different demographic groups. In this work, we measure…

Machine Learning · Computer Science 2021-06-01 Runshan Fu , Yangfan Liang , Peter Zhang

There has been a growing interest in covariate adjustment in the analysis of randomized controlled trials in past years. For instance, the U.S. Food and Drug Administration recently issued guidance that emphasizes the importance of…

Methodology · Statistics 2023-06-12 Kelly Van Lancker , Frank Bretz , Oliver Dukes

In a stratified clinical trial design with time to event end points, stratification factors are often accounted for the log-rank test and the Cox regression analyses. In this work, we have evaluated the impact of inclusion of stratification…

Applications · Statistics 2020-06-30 Madan G. Kundu , Shoubhik Mondal

Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging…

Motivated by a potential-outcomes perspective, the idea of principal stratification has been widely recognized for its relevance in settings susceptible to posttreatment selection bias such as randomized clinical trials where treatment…

Applications · Statistics 2011-11-08 Corwin M. Zigler , Thomas R. Belin

The incorporation of unlabeled data in regression and classification analysis is an increasing focus of the applied statistics and machine learning literatures, with a number of recent examples demonstrating the potential for unlabeled data…

Methodology · Statistics 2009-09-29 Feng Liang , Sayan Mukherjee , Mike West

For obtaining causal inferences that are objective, and therefore have the best chance of revealing scientific truths, carefully designed and executed randomized experiments are generally considered to be the gold standard. Observational…

Applications · Statistics 2008-11-12 Donald B. Rubin

In this review, we present econometric and statistical methods for analyzing randomized experiments. For basic experiments we stress randomization-based inference as opposed to sampling-based inference. In randomization-based inference,…

Methodology · Statistics 2017-10-26 Susan Athey , Guido Imbens

Over the past decades, researchers and ML practitioners have come up with better and better ways to build, understand and improve the quality of ML models, but mostly under the key assumption that the training data is distributed…

Machine Learning · Computer Science 2019-10-14 Yeounoh Chung , Peter J. Haas , Eli Upfal , Tim Kraska

Survival models incorporating random effects to account for unmeasured heterogeneity are being increasingly used in biostatistical and applied research. Specifically, unmeasured covariates whose lack of inclusion in the model would lead to…

Methodology · Statistics 2020-05-06 Alessandro Gasparini , Mark S. Clements , Keith R. Abrams , Michael J. Crowther

This paper considers the problem of inference in cluster randomized experiments when cluster sizes are non-ignorable. Here, by a cluster randomized experiment, we mean one in which treatment is assigned at the cluster level. By…

Econometrics · Economics 2024-04-11 Federico Bugni , Ivan Canay , Azeem Shaikh , Max Tabord-Meehan

In clinical and epidemiological research doubly truncated data often appear. This is the case, for instance, when the data registry is formed by interval sampling. Double truncation generally induces a sampling bias on the target variable,…

Methodology · Statistics 2023-01-11 Jacobo de Uña-Álvarez

Baseline covariates in randomized experiments are often used in the estimation of treatment effects, for example, when estimating treatment effects within covariate-defined subgroups. In practice, however, covariate values may be missing…

Methodology · Statistics 2023-08-29 Gauri Kamat , Jerome P. Reiter

Automated classifiers (ACs), often built via supervised machine learning (SML), can categorize large, statistically powerful samples of data ranging from text to images and video, and have become widely popular measurement devices in…

Machine Learning · Computer Science 2024-12-10 Nathan TeBlunthuis , Valerie Hase , Chung-Hong Chan

Statistical models that include random effects are commonly used to analyze longitudinal and correlated data, often with strong and parametric assumptions about the random effects distribution. There is marked disagreement in the literature…

Methodology · Statistics 2012-01-11 Charles E. McCulloch , John M. Neuhaus

An inductive probabilistic classification rule must generally obey the principles of Bayesian predictive inference, such that all observed and unobserved stochastic quantities are jointly modeled and the parameter uncertainty is fully…

Machine Learning · Statistics 2015-03-25 Henrik Nyman , Jie Xiong , Johan Pensar , Jukka Corander

In classification problems, sampling bias between training data and testing data is critical to the ranking performance of classification scores. Such bias can be both unintentionally introduced by data collection and intentionally…

Methodology · Statistics 2017-11-02 Chandler Zuo

An important phenomenon in high dimensional biological data is the presence of unobserved covariates that can have a significant impact on the measured response. When these factors are also correlated with the covariate(s) of interest (i.e.…

Methodology · Statistics 2018-02-02 Chris McKennan , Dan Nicolae