English
Related papers

Related papers: Learning under Singularity: An Information Criteri…

200 papers

We introduce a new criterion to determine the order of an autoregressive model fitted to time series data. It has the benefits of the two well-known model selection techniques, the Akaike information criterion and the Bayesian information…

Statistics Theory · Mathematics 2016-08-25 Jie Ding , Vahid Tarokh , Yuhong Yang

In this work, we introduce a new information-theoretic perspective on Multiple Instance Learning (MIL) for parameter estimation with i.i.d. data, and show that MIL can outperform single-instance learners in low-signal regimes. Prior work…

Machine Learning · Computer Science 2025-12-03 Atakan Azakli , Bernd Stelzer

In multivariate extreme value analysis, the estimation of the dependence structure in extremes is demanding, especially in the context of high-dimensional data. Therefore, a common approach is to reduce the model dimension by considering…

Methodology · Statistics 2025-07-08 Lucas Butsch , Vicky Fasen-Hartmann

The Information Bottleneck (IB) principle facilitates effective representation learning by preserving label-relevant information while compressing irrelevant information. However, its strong reliance on accurate labels makes it inherently…

Machine Learning · Computer Science 2025-12-12 Yi Huang , Qingyun Sun , Yisen Gao , Haonan Yuan , Xingcheng Fu , Jianxin Li

The Akaike information criterion (AIC) is a common tool for model selection. It is frequently used in violation of regularity conditions at parameter space singularities and boundaries. The expected AIC is generally not asymptotically…

Statistics Theory · Mathematics 2022-11-09 Jonathan D. Mitchell , Elizabeth S. Allman , John A. Rhodes

We consider class incremental learning (CIL) problem, in which a learning agent continuously learns new classes from incrementally arriving training data batches and aims to predict well on all the classes learned so far. The main challenge…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Hongjoon Ahn , Jihwan Kwak , Subin Lim , Hyeonsu Bang , Hyojun Kim , Taesup Moon

Watanabe-Akaike information criterion (WAIC; Watanabe, 2010) and leave-one-out cross validation (LOO) are two fully Bayesian model selection methods that have been shown to perform better than other traditional information-criterion based…

Applications · Statistics 2018-06-27 Luo Yong

This paper introduces and develops a theoretical extension of the widely applicable information criterion (WAIC), called the Covariance-Corrected WAIC (CC-WAIC), that applied for Bayesian sequential data models. The CC-WAIC accounts for…

Methodology · Statistics 2025-09-23 Safaa K. Kadhem

Many problems in statistics and machine learning can be formulated as model selection problems, where the goal is to choose an optimal parsimonious model among a set of candidate models. It is typical to conduct model selection by…

Methodology · Statistics 2024-04-29 Qingyuan Zhang , Hien Duy Nguyen

Singular statistical models-including mixtures, matrix factorization, and neural networks-violate regular asymptotics due to parameter non-identifiability and degenerate Fisher geometry. Although singular learning theory characterizes…

Machine Learning · Statistics 2026-03-06 Sean Plummer

Bayesian diagnostic classification models (Bayesian DCMs) are effective for diagnosing students' skills. Research on the evaluation of relative model fit indices for DCMs using Bayesian estimation, however, is deficient. This study…

Applications · Statistics 2024-10-07 Ae Kyong Jung , Jonathan Templin

We review the Akaike, deviance, and Watanabe-Akaike information criteria from a Bayesian perspective, where the goal is to estimate expected out-of-sample-prediction error using a biascorrected adjustment of within-sample error. We focus on…

Methodology · Statistics 2013-07-24 Andrew Gelman , Jessica Hwang , Aki Vehtari

Noting the erroneous proclivity of information-theoretic approaches, like the Akaike information criterion (AIC), to select simpler models while performing model selection with a small sample size, we address the problem of new physics…

High Energy Physics - Phenomenology · Physics 2020-08-12 Srimoy Bhattacharya , Soumitra Nandi , Sunando Kumar Patra , Shantanu Sahoo

Non-concave penalized maximum likelihood methods, such as the Bridge, the SCAD, and the MCP, are widely used because they not only do parameter estimation and variable selection simultaneously but also have a high efficiency as compared to…

Methodology · Statistics 2015-12-31 Yuta Umezu , Yusuke Shimizu , Hiroki Masuda , Yoshiyuki Ninomiya

Model evaluation is a critical component in supervised machine learning classification analyses. Traditional metrics do not currently incorporate case difficulty. This renders the classification results unbenchmarked for generalization.…

Machine Learning · Computer Science 2023-02-10 Adrienne Kline , Joon Lee

We investigate the HSIC (Hilbert-Schmidt independence criterion) bottleneck as a regularizer for learning an adversarially robust deep neural network classifier. In addition to the usual cross-entropy loss, we add regularization terms for…

Machine Learning · Computer Science 2021-10-27 Zifeng Wang , Tong Jian , Aria Masoomi , Stratis Ioannidis , Jennifer Dy

In-context learning (ICL) has emerged as a particularly remarkable characteristic of Large Language Models (LLM): given a pretrained LLM and an observed dataset, LLMs can make predictions for new data points from the same distribution…

Machine Learning · Statistics 2024-06-04 Fabian Falck , Ziyu Wang , Chris Holmes

Many learning machines that have hierarchical structure or hidden variables are now being used in information science, artificial intelligence, and bioinformatics. However, several learning machines used in such fields are not regular but…

Machine Learning · Computer Science 2015-05-13 Sumio Watanabe

Cosine similarity between two documents can be computed using token embeddings formed by Large Language Models (LLMs) such as GPT-4, and used to categorize those documents across a range of uses. However, these similarities are ultimately…

Computation and Language · Computer Science 2025-05-16 Tailia Malloy , Maria José Ferreira , Fei Fang , Cleotilde Gonzalez

We try to establish a unified information theoretic approach to learning and to explore some of its applications. First, we define {\em predictive information} as the mutual information between the past and the future of a time series,…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Ilya Nemenman