中文
相关论文

相关论文: The leave-one-covariate-out conditional randomizat…

200 篇论文

Model inference, such as model comparison, model checking, and model selection, is an important part of model development. Leave-one-out cross-validation (LOO) is a general approach for assessing the generalizability of a model, but…

机器学习 · 统计学 2020-08-12 Måns Magnusson , Michael Riis Andersen , Johan Jonasson , Aki Vehtari

Consider a case-control study in which we have a random sample, constructed in such a way that the proportion of cases in our sample is different from that in the general population---for instance, the sample is constructed to achieve a…

统计方法学 · 统计学 2019-01-01 Rina Foygel Barber , Emmanuel Candes

Approximate Leave-One-Out Cross-Validation (ALO-CV) is a method that has been proposed to estimate the generalization error of a regularized estimator in the high-dimensional regime where dimension and sample size are of the same order, the…

统计理论 · 数学 2026-02-13 Pierre C Bellec

Two key tasks in high-dimensional regularized regression are tuning the regularization strength for accurate predictions and estimating the out-of-sample risk. It is known that the standard approach -- $k$-fold cross-validation -- is…

统计理论 · 数学 2025-10-24 Kevin Luo , Yufan Li , Pragya Sur

Conditional independence (CI) is central to causal inference, feature selection, and graphical modeling, yet it is untestable in many settings without additional assumptions. Existing CI tests often rely on restrictive structural…

机器学习 · 计算机科学 2025-12-23 Alek Frohlich , Vladimir Kostic , Karim Lounici , Daniel Perazzo , Massimiliano Pontil

Modern machine learning models are highly expressive but notoriously difficult to analyze statistically. In particular, while black-box predictors can achieve strong empirical performance, they rarely provide valid hypothesis tests or…

机器学习 · 计算机科学 2026-03-10 Mohamed Salem

In recent decades, multilevel regression and poststratification (MRP) has surged in popularity for population inference. However, the validity of the estimates can depend on details of the model, and there is currently little research on…

统计方法学 · 统计学 2022-09-07 Swen Kuh , Lauren Kennedy , Qixuan Chen , Andrew Gelman

We propose the conditional predictive impact (CPI), a consistent and unbiased estimator of the association between one or several features and a given outcome, conditional on a reduced feature set. Building on the knockoff framework of…

统计方法学 · 统计学 2021-05-14 David S. Watson , Marvin N. Wright

Leave-one-problem-out (LOPO) performance prediction requires machine learning (ML) models to extrapolate algorithms' performance from a set of training problems to a previously unseen problem. LOPO is a very challenging task even for…

机器学习 · 计算机科学 2023-06-01 Ana Nikolikj , Michal Pluháček , Carola Doerr , Peter Korošec , Tome Eftimov

Statistical power is often a concern for clustered RCTs due to variance inflation from design effects and the high cost of adding study clusters (such as hospitals, schools, or communities). While covariate pre-specification is the…

统计方法学 · 统计学 2020-05-07 Peter Z. Schochet

We consider the problem of constructing multiple independent conditional randomization tests using a single dataset. Because the tests are independent, the randomization p-values can be interpreted individually and combined using standard…

统计理论 · 数学 2024-10-14 Yao Zhang , Qingyuan Zhao

Researchers in biomedical studies often work with samples that are not selected uniformly at random from the population of interest, a major example being a case-control study. While these designs are motivated by specific scientific…

Lasso and other regularization procedures are attractive methods for variable selection, subject to a proper choice of shrinkage parameter. Given a set of potential subsets produced by a regularization algorithm, a consistent model…

统计方法学 · 统计学 2014-02-26 Minh-Ngoc Tran

Cross-validation is a common method for estimating the predictive performance of machine learning models. In a data-scarce regime, where one typically wishes to maximize the number of instances used for training the model, an approach…

统计方法学 · 统计学 2025-03-25 George I. Austin , Itsik Pe'er , Tal Korem

Adjusting for (baseline) covariates with working regression models becomes standard practice in the analysis of randomized clinical trials (RCT). When the dimension $p$ of the covariates is large relative to the sample size $n$,…

统计方法学 · 统计学 2025-12-24 Yujia Gu , Lin Liu , Wei Ma

We analyze the performance of cross-validation (CV) in the density estimation framework with two purposes: (i) risk estimation and (ii) model selection. The main focus is given to the so-called leave-$p$-out CV procedure (Lpo), where $p$…

统计理论 · 数学 2014-10-02 Alain Celisse

A practical limitation of cluster randomized controlled trials (cRCTs) is that the number of available clusters may be small, resulting in an increased risk of baseline imbalance under simple randomization. Constrained randomization…

统计方法学 · 统计学 2022-01-19 Yunji Zhou , Elizabeth L. Turner , Ryan A. Simmons , Fan Li

When conducting a randomized controlled trial, it is common to specify in advance the statistical analyses that will be used to analyze the data. Typically these analyses will involve adjusting for small imbalances in baseline covariates.…

应用统计 · 统计学 2017-08-04 Edward Wu , Johann Gagnon-Bartsch

This paper studies inference for quadratic forms of linear regression coefficients with clustered data and many covariates. Our framework covers three important special cases: instrumental variables regression with many instruments and…

计量经济学 · 经济学 2026-02-18 Michal Kolesár , Pengjin Min , Wenjie Wang , Yichong Zhang

Testing conditional independence has many applications, such as in Bayesian network learning and causal discovery. Different test methods have been proposed. However, existing methods generally can not work when only discretized…

机器学习 · 统计学 2025-03-19 Boyang Sun , Yu Yao , Guang-Yuan Hao , Yumou Qiu , Kun Zhang