中文
相关论文

相关论文: A Comparative Study of Model Selection Criteria fo…

200 篇论文

We present a novel data-driven strategy to choose the hyperparameter $k$ in the $k$-NN regression estimator without using any hold-out data. We treat the problem of choosing the hyperparameter as an iterative procedure (over $k$) and…

机器学习 · 统计学 2024-07-18 Yaroslav Averyanov , Alain Celisse

Symbolic Regression tries to find a mathematical expression that describes the relationship of a set of explanatory variables to a measured variable. The main objective is to find a model that minimizes the error and, optionally, that also…

人工智能 · 计算机科学 2018-02-27 Fabricio Olivetti de Franca

A maximum likelihood based model selection of discrete Bayesian networks is considered. The model selection is performed through scoring function $S$, which, for a given network $G$ and $n$-sample $D_n$, is defined to be the maximum…

统计理论 · 数学 2013-04-18 Nikolay H. Balov

Model selection is a ubiquitous problem that arises in the application of many statistical and machine learning methods. In the likelihood and related settings, it is typical to use the method of information criteria (IC) to choose the most…

统计理论 · 数学 2024-08-13 Hien Duy Nguyen

This paper considers the problem of approximating a density when it can be evaluated up to a normalizing constant at a limited number of points. We call this problem the Boltzmann approximation (BA) problem. The BA problem is ubiquitous in…

统计方法学 · 统计学 2020-10-08 Youngjun Choe , Yen-Chi Chen , Nick Terry

Symbolic regression (SR) aims to discover the underlying mathematical expressions that explain observed data. This holds promise for both gaining scientific insight and for producing inherently interpretable and generalizable models for…

机器学习 · 计算机科学 2026-02-05 David Otte , Jörg K. H. Franke , Arbër Zela , Fábio Ferreira , Frank Hutter

Model selection is indispensable to high-dimensional sparse modeling in selecting the best set of covariates among a sequence of candidate models. Most existing work assumes implicitly that the model is correctly specified or of fixed…

统计理论 · 数学 2014-12-24 Pallavi Basu , Yang Feng , Jinchi Lv

Many problems in statistics and machine learning can be formulated as model selection problems, where the goal is to choose an optimal parsimonious model among a set of candidate models. It is typical to conduct model selection by…

统计方法学 · 统计学 2024-04-29 Qingyuan Zhang , Hien Duy Nguyen

We consider Bayesian model selection in generalized linear models that are high-dimensional, with the number of covariates p being large relative to the sample size n, but sparse in that the number of active covariates is small compared to…

统计理论 · 数学 2011-12-26 Rina Foygel , Mathias Drton

In Symbolic Regression (SR), Genetic Programming (GP) is a popular search algorithm that delivers state-of-the-art results in term of accuracy. Its success relies on the concept of neutrality, which induces large plateaus that the search…

机器学习 · 计算机科学 2025-11-04 Fabricio Olivetti de Franca , Gabriel Kronberger

Data-driven model discovery (DDMD) algorithms are powerful tools for extracting interpretable symbolic models from data. However, identifying the model that best balances goodness-of-fit and sparsity is often a laborious process requiring…

定量方法 · 定量生物学 2026-02-26 Michael C Chung , Alen Zacharia , Juan Guan

Popular statistical software provides Bayesian information criterion (BIC) for multilevel models or linear mixed models. However, it has been observed that the combination of statistical literature and software documentation has led to…

统计方法学 · 统计学 2022-06-24 Sun-Joo Cho , Hao Wu , Matthew Naveiras

The purpose of this article is to look at how information criteria, such as AIC and BIC, relate to the g%SD fit criterion derived in Waddell et al. (2007, 2010a). The g%SD criterion measures the fit of data to model based on a normalized…

基因组学 · 定量生物学 2013-01-01 Peter J. Waddell , Xi Tan

Approximate Bayesian computation (ABC) methods make use of comparisons between simulated and observed summary statistics to overcome the problem of computationally intractable likelihood functions. As the practical implementation of ABC…

统计方法学 · 统计学 2013-06-12 M. G. B. Blum , M. A. Nunes , D. Prangle , S. A. Sisson

Most of the regularization methods such as the LASSO have one (or more) regularization parameter(s), and to select the value of the regularization parameter is essentially equal to select a model. Thus, to obtain a model suitable for the…

统计方法学 · 统计学 2025-11-07 Sumito Kurata , Kei Hirose

Model selection and order selection problems frequently arise in statistical practice. A popular approach to addressing these problems in the frequentist setting involves information criteria based on penalised maxima of log-likelihoods for…

统计理论 · 数学 2025-10-29 Hien Duy Nguyen , Mayetri Gupta , Jacob Westerhout , TrungTin Nguyen

Symbolic regression (SR) is the problem of learning a symbolic expression from numerical data. Recently, deep neural models trained on procedurally-generated synthetic datasets showed competitive performance compared to more classical…

机器学习 · 计算机科学 2023-05-11 Pierre-Alexandre Kamienny , Guillaume Lample , Sylvain Lamprier , Marco Virgolin

Double-descent refers to the unexpected drop in test loss of a learning algorithm beyond an interpolating threshold with over-parameterization, which is not predicted by information criteria in their classical forms due to the limitations…

机器学习 · 计算机科学 2023-11-15 Haobo Chen , Yuheng Bu , Gregory W. Wornell

We study model selection by the Bayesian information criterion (BIC) in fixed-dimensional exploratory factor analysis over a fixed finite family of compact covariance classes. Our main result shows that the BIC is strongly consistent for…

统计理论 · 数学 2026-04-10 Hien Duy Nguyen , Kei Hirose

The Minimum Description Length (MDL) principle selects the model that has the shortest code for data plus model. We show that for a countable class of models, MDL predictions are close to the true distribution in a strong sense. The result…

概率论 · 数学 2010-12-30 Marcus Hutter