English
Related papers

Related papers: Fast leave-one-cluster-out cross-validation using …

200 papers

Leave-one-out cross-validation (LOO) and the widely applicable information criterion (WAIC) are methods for estimating pointwise out-of-sample prediction accuracy from a fitted Bayesian model using the log-likelihood evaluated at the…

Computation · Statistics 2017-12-18 Aki Vehtari , Andrew Gelman , Jonah Gabry

The widely applicable information criterion (WAIC) has been used as a model selection criterion for Bayesian statistics in recent years. It is an asymptotically unbiased estimator of the Kullback-Leibler divergence between a Bayesian…

Methodology · Statistics 2022-08-09 Yoshiyuki Ninomiya

Cluster analysis is a popular unsupervised learning tool used in many disciplines to identify heterogeneous sub-populations within a sample. However, validating cluster analysis results and determining the number of clusters in a data set…

Machine Learning · Statistics 2024-04-26 Ali Turfah , Xiaoquan Wen

Model selection in linear regression models is a major challenge when dealing with high-dimensional data where the number of available measurements (sample size) is much smaller than the dimension of the parameter space. Traditional methods…

Signal Processing · Electrical Eng. & Systems 2023-07-05 Prakash B. Gohain , Magnus Jansson

Performing model selection between Gibbs random fields is a very challenging task. Indeed, due to the Markovian dependence structure, the normalizing constant of the fields cannot be computed using standard analytical or numerical methods.…

Computation · Statistics 2019-09-04 Julien Stoehr , Jean-Michel Marin , Pierre Pudlo

The Bayesian and Akaike information criteria aim at finding a good balance between under- and over-fitting. They are extensively used every day by practitioners. Yet we contend they suffer from at least two afflictions: their penalty…

Statistics Theory · Mathematics 2026-03-20 Sylvain Sardy , Maxime van Cutsem , Sara van de Geer

Clustered data are common in biomedical research. Observations in the same cluster are often more similar to each other than to observations from other clusters. The intraclass correlation coefficient (ICC), first introduced by R. A.…

Methodology · Statistics 2024-02-20 Shengxin Tu , Chun Li , Donglin Zeng , Bryan E. Shepherd

We develop an algorithm for model selection which allows for the consideration of a combinatorially large number of candidate models governing a dynamical system. The innovation circumvents a disadvantage of standard model selection which…

Data Analysis, Statistics and Probability · Physics 2017-11-01 Niall M. Mangan , J. Nathan Kutz , Steven L. Brunton , Joshua L. Proctor

It has been shown that AIC-type criteria are asymptotically efficient selectors of the tuning parameter in non-concave penalized regression methods under the assumption that the population variance is known or that a consistent estimator is…

Machine Learning · Statistics 2017-03-02 Cheryl J. Flynn , Clifford M. Hurvich , Jeffrey S. Simonoff

Cluster analysis requires many decisions: the clustering method and the implied reference model, the number of clusters and, often, several hyper-parameters and algorithms' tunings. In practice, one produces several partitions, and a final…

Machine Learning · Statistics 2023-08-14 Luca Coraggio , Pietro Coretto

In the field of spatial data analysis, spatially varying coefficients (SVC) models, which allow regression coefficients to vary by region and flexibly capture spatial heterogeneity, have continued to be developed in various directions.…

Methodology · Statistics 2025-10-14 Yuko Kakikawa , Yoshiyuki Ninomiya

The Akaike information criterion (AIC) has been used as a statistical criterion to compare the appropriateness of different dark energy candidate models underlying a particular data set. Under suitable conditions, the AIC is an indirect…

Cosmology and Nongalactic Astrophysics · Physics 2015-05-28 Ming Yang Jeremy Tan , Rahul Biswas

There is no, nor will there ever be, single best clustering algorithm. Nevertheless, we would still like to be able to distinguish between methods that work well on certain task types and those that systematically underperform. Clustering…

Machine Learning · Computer Science 2025-10-16 Marek Gagolewski

Classical confidence intervals after best subset selection are widely implemented in statistical software and are routinely used to guide practitioners in scientific fields to conclude significance. However, there are increasing concerns in…

Methodology · Statistics 2023-11-27 Huiming Lin , Meng Li

Stochastic blockmodels and variants thereof are among the most widely used approaches to community detection for social networks and relational data. A stochastic blockmodel partitions the nodes of a network into disjoint sets, called…

Methodology · Statistics 2015-09-16 Diego Franco Saldana , Yi Yu , Yang Feng

We introduce a new criterion to determine the order of an autoregressive model fitted to time series data. It has the benefits of the two well-known model selection techniques, the Akaike information criterion and the Bayesian information…

Statistics Theory · Mathematics 2016-08-25 Jie Ding , Vahid Tarokh , Yuhong Yang

Variable selection is essential for improving inference and interpretation in multivariate linear regression. Although a number of alternative regressor selection criteria have been suggested, the most prominent and widely used are the…

Statistics Theory · Mathematics 2020-01-07 Zhidong Bai , Yasunori Fujikoshi , Jiang Hu

We have recently proposed a new information-based approach to model selection, the Frequentist Information Criterion (FIC), that reconciles information-based and frequentist inference. The purpose of this current paper is to provide a…

Data Analysis, Statistics and Probability · Physics 2015-06-23 Paul A. Wiggins

Model selection is the problem of distinguishing competing models, perhaps featuring different numbers of parameters. The statistics literature contains two distinct sets of tools, those based on information theory such as the Akaike…

Astrophysics · Physics 2014-10-13 Andrew R Liddle

Missing values in tabular data restrict the use and performance of machine learning, requiring the imputation of missing values. The most popular imputation algorithm is arguably multiple imputations using chains of equations (MICE), which…

Machine Learning · Computer Science 2022-03-01 Manar D Samad , Sakib Abrar , Norou Diawara