中文
相关论文

相关论文: Cross validation for model selection: a primer wit…

200 篇论文

Cross-validation is the de facto standard for predictive model evaluation and selection. In proper use, it provides an unbiased estimate of a model's predictive performance. However, data sets often undergo various forms of data-dependent…

统计方法学 · 统计学 2023-01-18 Amit Moscovich , Saharon Rosset

Cross-validation is a standard tool for obtaining a honest assessment of the performance of a prediction model. The commonly used version repeatedly splits data, trains the prediction model on the training set, evaluates the model…

机器学习 · 统计学 2025-10-10 Tianyu Pan , Vincent Z. Yu , Viswanath Devanarayan , Lu Tian

Linear mixed effects models are highly flexible in handling a broad range of data types and are therefore widely used in applications. A key part in the analysis of data is model selection, which often aims to choose a parsimonious model…

统计方法学 · 统计学 2013-06-12 Samuel Müller , J. L. Scealy , A. H. Welsh

Robust model-fitting to spectroscopic transitions is a requirement across many fields of science. The corrected Akaike and Bayesian information criteria (AICc and BIC) are most frequently used to select the optimal number of fitting…

天体物理仪器与方法 · 物理学 2020-11-25 John K. Webb , Chung-Chi Lee , Robert F. Carswell , Dinko Milaković

While the Bayesian Information Criterion (BIC) and Akaike Information Criterion (AIC) are powerful tools for model selection in linear regression, they are built on different prior assumptions and thereby apply to different data generation…

统计方法学 · 统计学 2017-12-15 MB de Kock , HC Eggers

This paper presents the first general (supervised) statistical learning framework for point processes in general spaces. Our approach is based on the combination of two new concepts, which we define in the paper: i) bivariate innovations,…

统计方法学 · 统计学 2021-03-03 Ottmar Cronie , Mehdi Moradi , Christophe A. N. Biscio

While many statistical models and methods are now available for network analysis, resampling network data remains a challenging problem. Cross-validation is a useful general tool for model selection and parameter tuning, but is not directly…

统计方法学 · 统计学 2020-05-04 Tianxi Li , Elizaveta Levina , Ji Zhu

Variable selection plays a fundamental role in high-dimensional data analysis. Various methods have been developed for variable selection in recent years. Well-known examples are forward stepwise regression (FSR) and least angle regression…

统计方法学 · 统计学 2018-02-01 Siliang Gong , Kai Zhang , Yufeng Liu

First, we analyze the variance of the Cross Validation (CV)-based estimators used for estimating the performance of classification rules. Second, we propose a novel estimator to estimate this variance using the Influence Function (IF)…

机器学习 · 统计学 2021-11-10 Waleed A. Yousef

Leave-one-out cross-validation (LOO) and the widely applicable information criterion (WAIC) are methods for estimating pointwise out-of-sample prediction accuracy from a fitted Bayesian model using the log-likelihood evaluated at the…

统计计算 · 统计学 2017-12-18 Aki Vehtari , Andrew Gelman , Jonah Gabry

Theoretical developments on cross validation (CV) have mainly focused on selecting one among a list of finite-dimensional models (e.g., subset or order selection in linear regression) or selecting a smoothing parameter (e.g., bandwidth for…

统计理论 · 数学 2008-12-18 Yuhong Yang

The use of Bayesian information criterion (BIC) in the model selection procedure is under the assumption that the observations are independent and identically distributed (i.i.d.). However, in practice, we do not always have i.i.d. samples.…

应用统计 · 统计学 2021-05-03 Nan Shen , Bárbara González

We analyze the performance of cross-validation (CV) in the density estimation framework with two purposes: (i) risk estimation and (ii) model selection. The main focus is given to the so-called leave-$p$-out CV procedure (Lpo), where $p$…

统计理论 · 数学 2014-10-02 Alain Celisse

This text is a survey on cross-validation. We define all classical cross-validation procedures, and we study their properties for two different goals: estimating the risk of a given estimator, and selecting the best estimator among a given…

统计理论 · 数学 2017-03-10 Sylvain Arlot

The effectiveness and validity of applying variation partitioning methods in community ecology has been questioned. Here, using mathematical deduction and numerical simulation, we made an attempt to uncover the underlying mechanisms…

种群与进化 · 定量生物学 2014-02-17 Youhua Chen

A bias correction to Akaike's information criterion (AIC) is derived for seemingly unrelated regressions models. The correction is of particular use when the sample size is not much larger than the number of fitted parameters. A…

统计方法学 · 统计学 2009-06-05 J. L. van Velsen

In many applications, we have access to the complete dataset but are only interested in the prediction of a particular region of predictor variables. A standard approach is to find the globally best modeling method from a set of candidate…

机器学习 · 统计学 2022-02-21 Jiawei Zhang , Jie Ding , Yuhong Yang

We develop an algorithm for model selection which allows for the consideration of a combinatorially large number of candidate models governing a dynamical system. The innovation circumvents a disadvantage of standard model selection which…

数据分析、统计与概率 · 物理学 2017-11-01 Niall M. Mangan , J. Nathan Kutz , Steven L. Brunton , Joshua L. Proctor

We consider prediction in multiple studies with potential differences in the relationships between predictors and outcomes. Our objective is to integrate data from multiple studies to develop prediction models for unseen studies. We propose…

统计方法学 · 统计学 2024-07-23 Boyu Ren , Prasad Patil , Francesca Dominici , Giovanni Parmigiani , Lorenzo Trippa

Cross-validation is a useful and generally applicable technique often employed in machine learning, including decision tree induction. An important disadvantage of straightforward implementation of the technique is its computational…

机器学习 · 计算机科学 2007-05-23 Hendrik Blockeel , Jan Struyf