中文
相关论文

相关论文: A note on adjusting $R^2$ for using with cross-val…

200 篇论文

In this article, we derive concentration inequalities for the cross-validation estimate of the generalization error for subagged estimators, both for classification and regressor. General loss functions and class of predictors with both…

机器学习 · 统计学 2010-11-24 Matthieu CORNEC

In many applications, we have access to the complete dataset but are only interested in the prediction of a particular region of predictor variables. A standard approach is to find the globally best modeling method from a set of candidate…

机器学习 · 统计学 2022-02-21 Jiawei Zhang , Jie Ding , Yuhong Yang

The present work aims at deriving theoretical guaranties on the behavior of some cross-validation procedures applied to the $k$-nearest neighbors ($k$NN) rule in the context of binary classification. Here we focus on the leave-$p$-out…

统计理论 · 数学 2017-10-13 Alain Celisse , Tristan Mary-Huard

Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the…

机器学习 · 统计学 2016-01-21 Simone Romano , Nguyen Xuan Vinh , James Bailey , Karin Verspoor

This paper studies V-fold cross-validation for model selection in least-squares density estimation. The goal is to provide theoretical grounds for choosing V in order to minimize the least-squares loss of the selected estimator. We first…

统计理论 · 数学 2015-10-13 Sylvain Arlot , Matthieu Lerasle

Composite likelihood inference has gained much popularity thanks to its computational manageability and its theoretical properties. Unfortunately, performing composite likelihood ratio tests is inconvenient because of their awkward…

统计计算 · 统计学 2014-08-01 Manuela Cattelan , Nicola Sartori

Using the LRT statistic, a model R^2 is proposed for the generalized linear mixed model for assessing the association between the correlated outcomes and fixed effects. The R^2 compares the full model to a null model with all fixed effects…

统计方法学 · 统计学 2010-11-03 Lloyd J. Edwards

Scholars frequently use covariate balance tests to test the validity of natural experiments and related designs. Unfortunately, when measured covariates are unrelated to potential outcomes, balance is uninformative about key identification…

统计方法学 · 统计学 2025-10-15 Clara Bicalho , Adam Bouyamourn , Thad Dunning

Two indicators are classically used to evaluate the quality of rule-based classification systems: predictive accuracy, i.e. the system's ability to successfully reproduce learning data and coverage, i.e. the proportion of possible cases for…

人工智能 · 计算机科学 2020-04-07 Nassim Dehouche

We consider the application of a popular penalised regression method, Ridge Regression, to data with very high dimensions and many more covariates than observations. Our motivation is the problem of out-of-sample prediction and the setting…

应用统计 · 统计学 2012-05-04 Erika Cule , Maria De Iorio

We study ridge estimation of the precision matrix in the high-dimensional setting where the number of variables is large relative to the sample size. We first review two archetypal ridge estimators and note that their utilized penalties do…

统计方法学 · 统计学 2016-06-17 Wessel N. van Wieringen , Carel F. W. Peeters

In traditional k-fold cross-validation, each instance is used ($k-1$) times for training and once for testing, leading to redundancy that lets many instances disproportionately influence the learning phase. We introduce Irredundant $k$-fold…

机器学习 · 计算机科学 2025-08-29 Jesus S. Aguilar-Ruiz

Recent years have seen substantial advances in our understanding of high-dimensional ridge regression, but existing theories assume that training examples are independent. By leveraging techniques from random matrix theory and free…

机器学习 · 统计学 2025-11-06 Alexander Atanasov , Jacob A. Zavatone-Veth , Cengiz Pehlevan

Covariate adjustment can improve precision in analyzing randomized experiments. With fully observed data, regression adjustment and propensity score weighting are asymptotically equivalent in improving efficiency over unadjusted analysis.…

统计方法学 · 统计学 2024-03-06 Anqi Zhao , Peng Ding , Fan Li

We propose a method for constructing p-values for general hypotheses in a high-dimensional linear model. The hypotheses can be local for testing a single regression parameter or they may be more global involving several up to all…

统计方法学 · 统计学 2013-10-14 Peter Bühlmann

Prediction error is critical to assessing the performance of statistical methods and selecting statistical models. We propose the cross-validation and approximated cross-validation methods for estimating prediction error under a broad…

统计理论 · 数学 2007-06-13 Jianqing Fan , Chunming Zhang

The estimated accuracy of a classifier is a random quantity with variability. A common practice in supervised machine learning, is thus to test if the estimated accuracy is significantly better than chance level. This method of signal…

统计方法学 · 统计学 2020-01-28 Jonathan D. Rosenblatt , Yuval Benjamini , Roee Gilron , Roy Mukamel , Jelle J. Goeman

Enriching existing predictive models with new biomolecular markers is an important task in the new multi-omic era. Clinical studies increasingly include new sets of omic measurements which may prove their added value in terms of predictive…

There has been a growing interest in covariate adjustment in the analysis of randomized controlled trials in past years. For instance, the U.S. Food and Drug Administration recently issued guidance that emphasizes the importance of…

统计方法学 · 统计学 2023-06-12 Kelly Van Lancker , Frank Bretz , Oliver Dukes

The coefficient of determination, the $R^2$, is often used to measure the variance explained by an affine combination of multiple explanatory covariates. An attribution of this explanatory contribution to each of the individual covariates…

统计方法学 · 统计学 2020-08-11 Daniel Fryer , Inga Strumke , Hien Nguyen