中文
相关论文

相关论文: Fast leave-one-cluster-out cross-validation using …

200 篇论文

In recent several years, the information bottleneck (IB) principle provides an information-theoretic framework for deep multi-view clustering (MVC) by compressing multi-view observations while preserving the relevant information of multiple…

信息论 · 计算机科学 2024-03-26 Xiaoqiang Yan , Zhixiang Jin , Fengshou Han , Yangdong Ye

It is crucial to assess the predictive performance of a model to establish its practicality and relevance in real-world scenarios, particularly for high-dimensional data analysis. Among data splitting or resampling methods, cross-validation…

统计方法学 · 统计学 2025-11-26 Iris Ivy Gauran , Hernando Ombao , Zhaoxia Yu

Clustering is a popular machine learning technique for data mining that can process and analyze datasets to automatically reveal sample distribution patterns. Since the ubiquitous categorical data naturally lack a well-defined metric space…

机器学习 · 计算机科学 2025-09-01 Yiqun Zhang , Mingjie Zhao , Hong Jia , Yang Lu , Mengke Li , Yiu-ming Cheung

There are several methods for model selection in cosmology which have at least two major goals, that of finding the correct model or predicting well. In this work we discuss through a study of well-known model selection methods like Akaike…

宇宙学与河外天体物理 · 物理学 2021-02-23 Mehdi Rezaei , Mohammad Malekjani

Model selection in clustering requires (i) to specify a suitable clustering principle and (ii) to control the model order complexity by choosing an appropriate number of clusters depending on the noise level in the data. We advocate an…

信息论 · 计算机科学 2010-06-03 Joachim M. Buhmann

For community detection problem, spectral clustering is a widely used method for detecting clusters in networks. In this paper, we propose an improved spectral clustering (ISC) approach under the degree corrected stochastic block model…

机器学习 · 统计学 2020-11-13 Huan Qing , Jingli Wang

A first step when fitting multilevel models to continuous responses is to explore the degree of clustering in the data. Researchers fit variance-component models and then report the proportion of variation in the response that is due to…

统计方法学 · 统计学 2020-02-17 George Leckie , William Browne , Harvey Goldstein , Juan Merlo , Peter Austin

A popular model selection approach for generalized linear mixed-effects models is the Akaike information criterion, or AIC. Among others, \cite{vaida05} pointed out the distinction between the marginal and conditional inference depending on…

统计方法学 · 统计学 2008-10-14 Heng Lian

Data clustering involves identifying latent similarities within a dataset and organizing them into clusters or groups. The outcomes of various clustering algorithms differ as they are susceptible to the intrinsic characteristics of the…

机器学习 · 计算机科学 2024-07-31 Bryar A. Hassan , Noor Bahjat Tayfor , Alla A. Hassan , Aram M. Ahmed , Tarik A. Rashid , Naz N. Abdalla

We describe a fast computation method for leave-one-out cross-validation (LOOCV) for $k$-nearest neighbours ($k$-NN) regression. We show that, under a tie-breaking condition for nearest neighbours, the LOOCV estimate of the mean square…

机器学习 · 统计学 2024-12-05 Motonobu Kanagawa

In data containing heterogeneous subpopulations, classification performance benefits from incorporating the knowledge of cluster structure in the classifier. Previous methods for such combined clustering and classification either 1) are…

机器学习 · 计算机科学 2023-01-04 Shivin Srivastava , Siddharth Bhatia , Lingxiao Huang , Lim Jun Heng , Kenji Kawaguchi , Vaibhav Rajan

Pattern discovery in multidimensional data sets has been the subject of research for decades. There exists a wide spectrum of clustering algorithms that can be used for this purpose. However, their practical applications share a common…

人工智能 · 计算机科学 2022-11-28 Szymon Bobek , Michał Kuk , Jakub Brzegowski , Edyta Brzychczy , Grzegorz J. Nalepa

Clustering is part of unsupervised analysis methods that consist in grouping samples into homogeneous and separate subgroups of observations also called clusters. To interpret the clusters, statistical hypothesis testing is often used to…

统计方法学 · 统计学 2022-10-25 Benjamin Hivert , Denis Agniel , Rodolphe Thiébaut , Boris P Hejblum

Cross-validation is a standard tool for obtaining a honest assessment of the performance of a prediction model. The commonly used version repeatedly splits data, trains the prediction model on the training set, evaluates the model…

机器学习 · 统计学 2025-10-10 Tianyu Pan , Vincent Z. Yu , Viswanath Devanarayan , Lu Tian

Credit risk default prediction remains a cornerstone of risk management in the financial industry. The task involves estimating the likelihood that a borrower will fail to meet debt obligations, an objective critical for lending decisions,…

机器学习 · 计算机科学 2026-04-21 Swattik Maiti , Ritik Pratap Singh , Fardina Fathmiul Alam

Interrelated Two-way Clustering (ITC) is an unsupervised clustering method developed to divide samples into two groups in gene expression data obtained through microarrays, selecting important genes simultaneously in the process. This has…

统计计算 · 统计学 2018-05-08 Subhabrata Majumdar , Subhash C. Basak , Gregory D. Grunwald

Generalization measures have been studied extensively in the machine learning community to better characterize generalization gaps. However, establishing a reliable generalization measure for statistically singular models such as deep…

机器学习 · 计算机科学 2026-02-27 Hiroki Naganuma , Taiji Suzuki , Rio Yokota , Masahiro Nomura , Kohta Ishikawa , Ikuro Sato

We propose a novel methodology for feature screening in clustering massive datasets, in which both the number of features and the number of observations can potentially be very large. Taking advantage of a fusion penalization based convex…

统计方法学 · 统计学 2017-10-05 Trambak Banerjee , Gourab Mukherjee , Peter Radchenko

Geometric Akaike Information Criteria (G-AICs) for generalized noise-level dependent crystallographic symmetry classifications of two-dimensional (2D) images that are more or less periodic in either two or one dimensions as well as Akaike…

应用物理 · 物理学 2018-01-08 Peter Moeck

Information criteria, such as Akaike's information criterion and Bayesian information criterion are often applied in model selection. However, their asymptotic behaviors for selecting geostatistical regression models have not been well…

统计理论 · 数学 2014-12-03 Chih-Hao Chang , Hsin-Cheng Huang , Ching-Kang Ing