English
Related papers

Related papers: A Comparative Study of Model Selection Criteria fo…

200 papers

Prior design is one of the most important problems in both statistics and machine learning. The cross validation (CV) and the widely applicable information criterion (WAIC) are predictive measures of the Bayesian estimation, however, it has…

Machine Learning · Computer Science 2015-03-30 Sumio Watanabe

Information bottleneck is an information-theoretic principle of representation learning that aims to learn a maximally compressed representation that preserves as much information about labels as possible. Under this principle, two…

Information Theory · Computer Science 2023-11-08 Yuyan Ni , Yanyan Lan , Ao Liu , Zhiming Ma

Bayesian diagnostic classification models (Bayesian DCMs) are effective for diagnosing students' skills. Research on the evaluation of relative model fit indices for DCMs using Bayesian estimation, however, is deficient. This study…

Applications · Statistics 2024-10-07 Ae Kyong Jung , Jonathan Templin

Traditional probabilistic methods for the simulation of advection-diffusion equations (ADEs) often overlook the entropic contribution of the discretization, e.g., the number of particles, within associated numerical methods. Many times, the…

Numerical Analysis · Mathematics 2021-06-15 Nhat Thanh Tran , David A. Benson , Michael J. Schmidt , Stephen D. Pankavich

Symbolic regression seeks to uncover physical laws from experimental data by searching for closed-form expressions, which is an important task in AI-driven scientific discovery. Yet the exponential growth of the search space of expression…

Symbolic Computation · Computer Science 2026-02-13 Nan Jiang , Ziyi Wang , Yexiang Xue

In statistical learning, models are classified as regular or singular depending on whether the mapping from parameters to probability distributions is injective. Most models with hierarchical structures or latent variables are singular, for…

Machine Learning · Statistics 2025-11-26 Naoki Hayashi , Takuro Kutsuna , Sawa Takamuku

Model selection in machine learning (ML) is a crucial part of the Bayesian learning procedure. Model choice may impose strong biases on the resulting predictions, which can hinder the performance of methods such as Bayesian neural networks…

Machine Learning · Statistics 2022-08-09 Simón Rodríguez Santana , Luis A. Ortega , Daniel Hernández-Lobato , Bryan Zaldívar

Model selection in mixed models based on the conditional distribution is appropriate for many practical applications and has been a focus of recent statistical research. In this paper we introduce the R-package cAIC4 that allows for the…

Computation · Statistics 2018-03-20 Benjamin Säfken , David Rügamer , Thomas Kneib , Sonja Greven

Performing model selection between Gibbs random fields is a very challenging task. Indeed, due to the Markovian dependence structure, the normalizing constant of the fields cannot be computed using standard analytical or numerical methods.…

Computation · Statistics 2019-09-04 Julien Stoehr , Jean-Michel Marin , Pierre Pudlo

Determining how to appropriately select the tuning parameter is essential in penalized likelihood methods for high-dimensional data analysis. We examine this problem in the setting of penalized likelihood methods for generalized linear…

Methodology · Statistics 2016-05-12 Yingying Fan , Cheng Yong Tang

Bayesian model averaging, model selection and its approximations such as BIC are generally statistically consistent, but sometimes achieve slower rates og convergence than other methods such as AIC and leave-one-out cross-validation. On the…

Statistics Theory · Mathematics 2008-09-17 Tim van Erven , Peter Grunwald , Steven de Rooij

Consider $n$ independent and identically distributed $p$-dimensional Gaussian random vectors with covariance matrix $\Sigma.$ The problem of estimating $\Sigma$ when $p$ is much larger than $n$ has received a lot of attention in recent…

Statistics Theory · Mathematics 2016-03-07 Danning Li , Hui Zou

In a regression task, a function is learned from labeled data to predict the labels at new data points. The goal is to achieve small prediction errors. In symbolic regression, the goal is more ambitious, namely, to learn an interpretable…

Machine Learning · Computer Science 2025-06-25 Paul Kahlmeyer , Joachim Giesen , Michael Habeck , Henrik Voigt

Smoothed AIC (S-AIC) and Smoothed BIC (S-BIC) are very widely used in model averaging and are very easily to implement. Especially, the optimal model averaging method MMA and JMA have only been well developed in linear models. Only by…

Methodology · Statistics 2019-10-29 Miaomiao Wang , Xinyu Zhang , Alan T. K. Wan , Guohua Zou

Mixture model-based clustering has become an increasingly popular data analysis technique since its introduction over fifty years ago, and is now commonly utilized within a family setting. Families of mixture models arise when the component…

Methodology · Statistics 2019-11-11 Sanjeena Subedi , Paul D. McNicholas

Choosing models from a well-fitted evolved population that generalizes beyond training data is difficult. We introduce a pragmatic method to estimate model complexity using Hessian rank for post-processing selection. Complexity is…

Machine Learning · Computer Science 2025-01-30 Nathan Haut , Zenas Huang , Adam Alessio

A vast amount of ecological knowledge generated recently has hinged upon the ability of model selection methods to discriminate among various ecological hypotheses. The last decade has seen the rise of Bayesian hierarchical models in…

Applications · Statistics 2018-10-08 Soumen Dey , Mohan Delampady , Arjun M. Gopalaswamy

A stochastic search method, the so-called Adaptive Subspace (AdaSub) method, is proposed for variable selection in high-dimensional linear regression models. The method aims at finding the best model with respect to a certain model…

Computation · Statistics 2021-04-20 Christian Staerk , Maria Kateri , Ioannis Ntzoufras

This paper introduces Kernel-based Information Criterion (KIC) for model selection in regression analysis. The novel kernel-based complexity measure in KIC efficiently computes the interdependency between parameters of the model using a…

Machine Learning · Statistics 2014-12-16 Somayeh Danafar , Kenji Fukumizu , Faustino Gomez

This work studies an information-theoretic performance limit of an integrated sensing and communication (ISAC) system where the goal of sensing is to estimate a random continuous state. Considering the mean-squared error (MSE) for…

Information Theory · Computer Science 2025-01-30 Daewon Seo