English
Related papers

Related papers: Overlap violations in external validity

200 papers

We present two sets of theoretical results on the grouped lasso with overlap of Jacob, Obozinski and Vert (2009) in the linear regression setting. This method allows for joint selection of predictors in sparse regression, allowing for…

Machine Learning · Statistics 2011-11-11 Daniel Percival

Many popular methods for building confidence intervals on causal effects under high-dimensional confounding require strong "ultra-sparsity" assumptions that may be difficult to validate in practice. To alleviate this difficulty, we here…

Statistics Theory · Mathematics 2019-05-06 Jelena Bradic , Stefan Wager , Yinchu Zhu

We consider the common setting where one observes probability estimates for a large number of events, such as default risks for numerous bonds. Unfortunately, even with unbiased estimates, selecting events corresponding to the most extreme…

Methodology · Statistics 2021-10-14 Gareth M. James , Peter Radchenko , Bradley Rava

Outcome-dependent sampling designs are common in many different scientific fields including epidemiology, ecology, and economics. As with all observational studies, such designs often suffer from unmeasured confounding, which generally…

Methodology · Statistics 2020-10-13 Erin E. Gabriel , Michael C. Sachs , Arvid Sjölander

External controls from historical trials or observational data can augment randomized controlled trials when large-scale randomization is impractical or unethical, such as in drug evaluation for rare diseases. However, non-randomized…

Methodology · Statistics 2025-05-08 Ke Zhu , Shu Yang , Xiaofei Wang

The ability to quantify distinctness of a cluster structure is fundamental for certain simulation studies, in particular for those comparing performance of different classification algorithms. The intrinsic integral measure based on the…

Statistics Theory · Mathematics 2014-07-29 Ewa Nowakowska , Jacek Koronacki , Stan Lipovetsky

Stress testing poses a causal question: how would portfolio credit losses change if the macroeconomy followed an adverse counterfactual path? Yet standard practice remains predictive and might be therefore vulnerable to omitted-variable…

Artificial Intelligence · Computer Science 2026-05-19 Yu Wang , Xiangchen Liu , Siguang Li

Randomized trials are widely considered as the gold standard for evaluating the effects of decision policies. Trial data is, however, drawn from a population which may differ from the intended target population and this raises a problem of…

Methodology · Statistics 2024-10-30 Sofia Ek , Dave Zachariah

Over the past two decades, considerable strides have been made in advancing neuroscientific techniques, yet challenges remain in attributing causality to observed associations. This review addresses a fundamental issue in observational…

Other Quantitative Biology · Quantitative Biology 2025-11-04 Eric W. Bridgeford , Brian S. Caffo , Maya B. Mathur , Russell A. Poldrack

Effective human-machine collaboration requires machine learning models to externalize uncertainty, so users can reflect and intervene when necessary. For language models, these representations of uncertainty may be impacted by sycophancy…

Computation and Language · Computer Science 2024-10-22 Anthony Sicilia , Mert Inan , Malihe Alikhani

While observational data are routinely used to estimate causal effects of biomedical treatments, doing so requires special methods to adjust for observed confounding. These methods invariably rely on untestable statistical and causal…

Methodology · Statistics 2026-03-02 Arman Oganisian

In most real-world systems units are interconnected and can be represented as networks consisting of nodes and edges. For instance, in social systems individuals can have social ties, family or financial relationships. In settings where…

Methodology · Statistics 2018-07-31 Laura Forastiere , Fabrizia Mealli , Albert Wu , Edoardo Airoldi

For testing the statistical significance of a treatment effect, we usually compare between two parts of a population, one is exposed to the treatment, and the other is not exposed to it. Standard parametric and nonparametric two-sample…

Computation · Statistics 2012-11-02 Bikram Karmakar , Kumaresh Dhara , Kushal Kumar Dey , Analabha Basu , Anil Ghosh

Causal discovery can be a powerful tool for investigating causality when a system can be observed but is inaccessible to experiments in practice. Despite this, it is rarely used in any scientific or medical fields. One of the major hurdles…

Machine Learning · Statistics 2019-10-07 Erich Kummerfeld , Alexander Rix

Causal decomposition has provided a powerful tool to analyze health disparity problems, by assessing the proportion of disparity caused by each mediator. However, most of these methods lack \emph{policy implications}, as they fail to…

Methodology · Statistics 2023-02-21 Xinwei Sun , Xiangyu Zheng , Jim Weinstein

Hill's specificity criterion has been highly influential in biomedical and epidemiological research. However, it remains controversial and its application often relies on subjective and qualitative analysis without a comprehensive and…

Methodology · Statistics 2025-06-24 Wang Miao

Propensity score (PS) weighting methods are often used in non-randomized studies to adjust for confounding and assess treatment effects. The most popular among them, the inverse probability weighting (IPW), assigns weights that are…

Methodology · Statistics 2020-11-04 Yunji Zhou , Roland A. Matsouaka , Laine Thomas

Assessing the fairness of a decision making system with respect to a protected class, such as gender or race, is challenging when class membership labels are unavailable. Probabilistic models for predicting the protected class based on…

Applications · Statistics 2018-11-28 Jiahao Chen , Nathan Kallus , Xiaojie Mao , Geoffry Svacha , Madeleine Udell

The conditional average treatment effect (CATE) is widely used in personalized medicine to inform therapeutic decisions. However, state-of-the-art methods for CATE estimation (so-called meta-learners) often perform poorly in the presence of…

Machine Learning · Computer Science 2026-03-12 Valentyn Melnychuk , Dennis Frauen , Jonas Schweisthal , Stefan Feuerriegel

Studies intended to estimate the effect of a treatment, like randomized trials, may not be sampled from the desired target population. To correct for this discrepancy, estimates can be transported to the target population. Methods for…