中文
相关论文

相关论文: Adjusting the adjusted Rand Index -- A multinomial…

200 篇论文

The adjusted Rand index (ARI) is commonly used in cluster analysis to measure the degree of agreement between two data partitions. Since its introduction, exploring the situations of extreme agreement and disagreement under different…

机器学习 · 统计学 2020-12-10 José E. Chacón , Ana I. Rastrojo

Adjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on Shannon information…

机器学习 · 统计学 2015-12-07 Simone Romano , Nguyen Xuan Vinh , James Bailey , Karin Verspoor

We consider the simultaneous clustering of rows and columns of a matrix and more particularly the ability to measure the agreement between two co-clustering partitions. The new criterion we developed is based on the Adjusted Rand Index and…

应用统计 · 统计学 2020-12-16 Valerie Robert , Yann Vasseur , Vincent Brault

The misclassification error distance and the adjusted Rand index are two of the most commonly used criteria to evaluate the performance of clustering algorithms. This paper provides an in-depth comparison of the two criteria, aimed to…

机器学习 · 统计学 2019-07-29 José E. Chacón

In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters…

机器学习 · 统计学 2019-01-08 Matthijs J. Warrens , Hanneke van der Hoef

Clustering is at the very core of machine learning, and its applications proliferate with the increasing availability of data. However, as datasets grow, comparing clusterings with an adjustment for chance becomes computationally difficult,…

机器学习 · 计算机科学 2023-08-01 Kai Klede , Leo Schwinn , Dario Zanca , Björn Eskofier

A well-known metric for quantifying the similarity between two clusterings is the adjusted mutual information. Compared to mutual information, a corrective term based on random permutations of the labels is introduced, preventing two…

机器学习 · 计算机科学 2021-03-24 Denys Lazarenko , Thomas Bonald

Focusing on the most significant features of a dataset is useful both in machine learning (ML) and data mining. In ML, it can lead to a higher accuracy, a faster learning process, and ultimately a simpler and more understandable model. In…

机器学习 · 计算机科学 2023-01-12 Suryani Lim , Henri Prade , Gilles Richard

The continuous net reclassification improvement (NRI) statistic is a popular model change measure that was developed to assess the incremental value of new factors in a risk prediction model. Two prominent statistical issues identified in…

统计方法学 · 统计学 2022-04-08 Glenn Heller

Adjusted similarity measures, such as Cohen's kappa for inter-rater reliability and the adjusted Rand index used to compare clustering algorithms, are a vital tool for comparing discrete labellings. These measures are intended to have the…

统计方法学 · 统计学 2026-01-16 William L. Lippitt , Edward J. Bedrick , Nichole E. Carlson

All Resolutions Inference (ARI) is a post hoc inference method for functional Magnetic Resonance Imaging (fMRI) data analysis that provides valid lower bounds on the proportion of truly active voxels within any, possibly data-driven,…

统计理论 · 数学 2025-11-05 Nils Peyrouset , Pierre Neuvial , Bertrand Thirion

Rerandomization is a strategy of increasing efficiency as compared to complete randomization. The idea with rerandomization is that of removing allocations with imbalance in the observed covariates and then randomizing within the set of…

统计方法学 · 统计学 2019-11-07 Junni L. Zhang , Per Johansson

Statistical approaches that successfully combine multiple datasets are more powerful, efficient, and scientifically informative than separate analyses. To address variation architectures correctly and comprehensively for high-dimensional…

统计方法学 · 统计学 2023-09-01 Jiuzhou Wang , Eric F. Lock

We provide a more efficient algorithm for computing the Rand Index when the data cluster comes from a change-point detection problem. Given $N$ data points and two clusterings of size $r$ and $s$, the algorithm runs on $O(r+s)$ time…

数据结构与算法 · 计算机科学 2025-06-23 Lucas de Oliveira Prates

Adaptive interventions (AIs) are increasingly becoming popular in medical and behavioral sciences. An AI is a sequence of individualized intervention options that specify for whom and under what conditions different intervention options…

应用统计 · 统计学 2018-12-18 Palash Ghosh , Inbal Nahum-Shani , Bonnie Spring , Bibhas Chakraborty

Random Indexing (RI) K-tree is the combination of two algorithms for clustering. Many large scale problems exist in document clustering. RI K-tree scales well with large inputs due to its low complexity. It also exhibits features that are…

信息检索 · 计算机科学 2010-02-02 Christopher M. De Vries , Lance De Vine , Shlomo Geva

Comparing clusterings is central to evaluating unsupervised models, yet the many existing similarity measures can produce widely divergent, sometimes contradictory, evaluations. Clustering similarity measures are typically organized into…

机器学习 · 统计学 2025-11-06 Alexander J. Gates

This paper investigates the application of consensus clustering and meta-clustering to the set of all possible partitions of a data set. We show that when using a "complement" of Rand Index as a measure of cluster similarity, the…

人工智能 · 计算机科学 2017-02-14 Mieczysław Kłopotek

Background: Imagine a paper with n nodes on it where each pair undergoes a coin toss experiment; if heads we connect the pair with an undirected link, while tails maintain the disconnection. This procedure yields a random graph. Now…

社会与信息网络 · 计算机科学 2023-12-29 Georgios Argyris

A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a…

声音 · 计算机科学 2020-05-21 Ricard Marxer , Hendrik Purwins
‹ 上一页 1 2 3 10 下一页 ›