中文
相关论文

相关论文: Fast leave-one-cluster-out cross-validation using …

200 篇论文

As the size $n$ of datasets become massive, many commonly-used clustering algorithms (for example, $k$-means or hierarchical agglomerative clustering (HAC) require prohibitive computational cost and memory. In this paper, we propose a…

Unsupervised anomaly detection (AD) is a fundamental problem in machine learning and statistics. A popular approach to unsupervised AD is clustering-based detection. However, this method lacks the ability to guarantee the reliability of the…

机器学习 · 统计学 2025-04-29 Nguyen Thi Minh Phu , Duong Tan Loc , Vo Nguyen Le Duy

Many modern data analyses benefit from explicitly modeling dependence structure in data -- such as measurements across time or space, ordered words in a sentence, or genes in a genome. A gold standard evaluation technique is structured…

Clustering techniques are often validated using benchmark datasets where class labels are used as ground-truth clusters. However, depending on the datasets, class labels may not align with the actual data clusters, and such misalignment…

机器学习 · 计算机科学 2025-03-04 Hyeon Jeon , Michaël Aupetit , DongHwa Shin , Aeri Cho , Seokhyeon Park , Jinwook Seo

Reporting test-retest reliability using the intraclass correlation coefficient (ICC) has received increasing attention due to the criticisms of poor transparency and replicability in neuroimaging research, as well as many other biomedical…

统计方法学 · 统计学 2026-01-07 Yufeng Liu , Xiangfei Hong , Shanbao Tong

This paper considers the problem of approximating a density when it can be evaluated up to a normalizing constant at a limited number of points. We call this problem the Boltzmann approximation (BA) problem. The BA problem is ubiquitous in…

统计方法学 · 统计学 2020-10-08 Youngjun Choe , Yen-Chi Chen , Nick Terry

We introduce a novel cross-validation method that we call latinCV and we compare this method to other model selection methods using data generated from a stochastic block model. Comparing latinCV to other cross-validation methods, we show…

统计方法学 · 统计学 2016-05-11 Beau Dabbs , Brian Junker

We present a model selection framework for the extraction of the CKM matrix element $|V_{cb}|$ from exclusive $B \to D^* l \nu$ decays. By framing the truncation of the Boyd-Grinstein-Lebed (BGL) parameterization as a model selection task,…

高能物理 - 唯象学 · 物理学 2024-12-11 Eric Persson , Florian Bernlochner

The misclassification error distance and the adjusted Rand index are two of the most commonly used criteria to evaluate the performance of clustering algorithms. This paper provides an in-depth comparison of the two criteria, aimed to…

机器学习 · 统计学 2019-07-29 José E. Chacón

Classical clustering methods do not provide users with direct control of the clustering results, and the clustering results may not be consistent with the relevant criterion that a user has in mind. In this work, we present a new…

计算机视觉与模式识别 · 计算机科学 2024-02-23 Sehyun Kwon , Jaeseung Park , Minkyu Kim , Jaewoong Cho , Ernest K. Ryu , Kangwook Lee

Classification tasks are widely investigated in the In-Context Learning (ICL) paradigm. However, current efforts are evaluated on disjoint benchmarks and settings, while their performances are significantly influenced by some trivial…

计算与语言 · 计算机科学 2025-04-21 Hakaze Cho , Naoya Inoue

Complex and larger networks are becoming increasingly prevalent in scientific applications in various domains. Although a number of models and methods exist for such networks, cross-validation on networks remains challenging due to the…

统计方法学 · 统计学 2026-03-12 Sayan Chakrabarty , Srijan Sengupta , Yuguo Chen

Determining the correct number of clusters (CNC) is an important task in data clustering and has a critical effect on finalizing the partitioning results. K-means is one of the popular methods of clustering that requires CNC. Validity index…

统计理论 · 数学 2019-11-28 Soosan Beheshti , Edward Nidoy , Faizan Rahman

The conventional clustering algorithms mine static databases and generate a set of patterns in the form of clusters. Many real life databases keep growing incrementally. For such dynamic databases, the patterns extracted from the original…

数据库 · 计算机科学 2013-10-28 A. M. Sowjanya , M. Shashi

K-fold cross-validation (CV) with squared error loss is widely used for evaluating predictive models, especially when strong distributional assumptions cannot be taken. However, CV with squared error loss is not free from distributional…

统计方法学 · 统计学 2021-08-10 Assaf Rabinowicz , Saharon Rosset

For linear models with a diverging number of parameters, it has recently been shown that modified versions of Bayesian information criterion (BIC) can identify the true model consistently. However, in many cases there is little…

统计方法学 · 统计学 2011-07-26 Heng Lian

In this paper, we propose a semi-supervised clustering method, CEC-IB, that models data with a set of Gaussian distributions and that retrieves clusters based on a partial labeling provided by the user (partition-level side information). By…

机器学习 · 计算机科学 2017-11-15 Marek Śmieja , Bernhard C. Geiger

We consider the problem of fast time-series data clustering. Building on previous work modeling the correlation-based Hamiltonian of spin variables we present an updated fast non-expensive Agglomerative Likelihood Clustering algorithm…

计算金融 · 定量金融 2022-03-22 Lionel Yelibi , Tim Gebbie

We propose two approaches for selecting variables in latent class analysis (i.e.,mixture model assuming within component independence), which is the common model-based clustering method for mixed data. The first approach consists in…

统计计算 · 统计学 2017-03-08 Matthieu Marbac , Mohammed Sedki

We go through the process of crafting a robust and numerically stable online algorithm for the computation of the Watanabe-Akaike information criteria (WAIC). We implement this algorithm in the NIMBLE software. The implementation is…

统计计算 · 统计学 2021-06-28 Joshua E. Hug , Christopher J. Paciorek