中文
相关论文

相关论文: The leave-one-covariate-out conditional randomizat…

200 篇论文

Leave-one-out cross-validation (LOO-CV) is a popular method for estimating out-of-sample predictive accuracy. However, computing LOO-CV criteria can be computationally expensive due to the need to fit the model multiple times. In the…

统计计算 · 统计学 2023-09-28 Luca Silva , Giacomo Zanella

Conditional independence testing (CIT) is essential for reliable scientific discovery. It prevents spurious findings and enables controlled feature selection. Recent CIT methods have used machine learning (ML) models as surrogates of the…

统计理论 · 数学 2026-02-02 Angel Reyero-Lobo , Bertrand Thirion , Pierre Neuvial

The out-of-sample error (OO) is the main quantity of interest in risk estimation and model selection. Leave-one-out cross validation (LO) offers a (nearly) distribution-free yet computationally demanding approach to estimate OO. Recent…

统计理论 · 数学 2023-10-27 Arnab Auddy , Haolin Zou , Kamiar Rahnama Rad , Arian Maleki

In many research fields, researchers aim to identify significant associations between a set of explanatory variables and a response while controlling the FDR. The Knockoff filter has been recently proposed in the frequentist paradigm to…

统计方法学 · 统计学 2026-04-22 Lorenzo Focardi-Olmi , Anna Gottard , Michele Guindani , Marina Vannucci

We propose a new approach to falsify causal discovery algorithms without ground truth, which is based on testing the causal model on a pair of variables that has been dropped when learning the causal model. To this end, we use the…

机器学习 · 统计学 2024-11-11 Daniela Schkoda , Philipp Faller , Patrick Blöbaum , Dominik Janzing

Estimating out-of-sample risk for models trained on large high-dimensional datasets is an expensive but essential part of the machine learning process, enabling practitioners to optimally tune hyperparameters. Cross-validation (CV) serves…

统计理论 · 数学 2025-04-28 Parth Nobel , Daniel LeJeune , Emmanuel J. Candès

Conditional local independence is an asymmetric independence relation among continuous time stochastic processes. It describes whether the evolution of one process is directly influenced by another process given the histories of additional…

统计理论 · 数学 2024-02-26 Alexander Mangulad Christgau , Lasse Petersen , Niels Richard Hansen

With machine learning being a popular topic in current computational materials science literature, creating representations for compounds has become common place. These representations are rarely compared, as evaluating their performance -…

机器学习 · 计算机科学 2023-05-26 Samantha Durdy , Michael Gaultois , Vladimir Gusev , Danushka Bollegala , Matthew J. Rosseinsky

Feature selection and importance estimation in a model-agnostic setting is an ongoing challenge of significant interest. Wrapper methods are commonly used because they are typically model-agnostic, even though they are computationally…

机器学习 · 统计学 2025-08-21 Chenghui Zheng , Garvesh Raskutti

Generalized linear mixed models (GLMM) are commonly used to analyze clustered data, but when the number of clusters is small to moderate, standard statistical tests may produce elevated type I error rates. Small-sample corrections have been…

统计方法学 · 统计学 2023-11-07 Hongxiang Qiu , Andrea J. Cook , Jennifer F. Bobb

Recursive partitioning approaches producing tree-like models are a long standing staple of predictive modeling, in the last decade mostly as ``sub-learners'' within state of the art ensemble methods like Boosting and Random Forest. However,…

机器学习 · 统计学 2015-12-14 Amichai Painsky , Saharon Rosset

Testing whether a variable of interest affects the outcome is one of the most fundamental problem in statistics and is often the main scientific question of interest. To tackle this problem, the conditional randomization test (CRT) is…

统计方法学 · 统计学 2023-05-26 Dae Woong Ham , Jiaze Qiu

Many contemporary large-scale applications involve building interpretable models linking a large set of potential covariates to a response in a nonlinear fashion, such as when the response is binary. Although this modeling problem has been…

统计方法学 · 统计学 2017-12-13 Emmanuel Candes , Yingying Fan , Lucas Janson , Jinchi Lv

The Model-X knockoff procedure has recently emerged as a powerful approach for feature selection with statistical guarantees. The advantage of knockoff is that if we have a good model of the features X, then we can identify salient features…

机器学习 · 统计学 2019-05-30 Jaime Roquero Gimenez , James Zou

Model-X knockoffs is a general procedure that can leverage any feature importance measure to produce a variable selection algorithm, which discovers true effects while rigorously controlling the number or fraction of false positives.…

统计方法学 · 统计学 2020-12-07 Zhimei Ren , Yuting Wei , Emmanuel Candès

Cross-validation can be used to measure a model's predictive accuracy for the purpose of model comparison, averaging, or selection. Standard leave-one-out cross-validation (LOO-CV) requires that the observation model can be factorized into…

统计方法学 · 统计学 2021-06-21 Paul-Christian Bürkner , Jonah Gabry , Aki Vehtari

Leveraging the large body of work devoted in recent years to describe redundancy and synergy in multivariate interactions among random variables, we propose a novel approach to quantify cooperative effects in feature importance, one of the…

数据分析、统计与概率 · 物理学 2025-03-14 Marlis Ontivero-Ortega , Luca Faes , Jesus M Cortes , Daniele Marinazzo , Sebastiano Stramaglia

A new statistical procedure (Model-X \cite{candes2018}) has provided a way to identify important factors using any supervised learning method controlling for FDR. This line of research has shown great potential to expand the horizon of…

统计方法学 · 统计学 2018-10-01 Ying Liu , Cheng Zheng

Per-instance automated algorithm configuration and selection are gaining significant moments in evolutionary computation in recent years. Two crucial, sometimes implicit, ingredients for these automated machine learning (AutoML) methods are…

神经与进化计算 · 计算机科学 2023-01-25 Ana Nikolikj , Carola Doerr , Tome Eftimov

Covariate adjustment is a general method for improving precision when estimating treatment effects in randomized trials and is recommended by the FDA in its 2023 guidance when baseline variables are prognostic for the primary outcome. We…