English
Related papers

Related papers: Establishing Validity for Distance Functions and I…

200 papers

Subject classification schemes are foundational to the organization, evaluation, and navigation of scientific knowledge. While expert-curated systems like Scopus provide widely used taxonomies, they often suffer from coarse granularity,…

Digital Libraries · Computer Science 2025-12-30 Zhuoqi Lyu , Qing Ke

A distinctive feature of a clustered observational study is its multilevel or nested data structure arising from the assignment of treatment, in a non-random manner, to groups or clusters of units or individuals. Examples are ubiquitous in…

Methodology · Statistics 2016-05-02 Luke Keele , Jose R. Zubizarreta

Model selection in clustering requires (i) to specify a suitable clustering principle and (ii) to control the model order complexity by choosing an appropriate number of clusters depending on the noise level in the data. We advocate an…

Information Theory · Computer Science 2010-06-03 Joachim M. Buhmann

In this paper, we address the problem of novel class discovery (NCD), which aims to cluster novel classes by leveraging knowledge from disjoint known classes. While recent advances have made significant progress in this area, existing NCD…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Xinhang Wan , Jiyuan Liu , Qian Qu , Suyuan Liu , Chuyu Zhang , Fangdi Wang , Xinwang Liu , En Zhu , Kunlun He

The LLM community often reports benchmark results as if they are synonymous with general model capabilities. However, benchmarks can have problems that distort performance, like test set contamination and annotator error. How can we know…

Artificial Intelligence · Computer Science 2026-02-18 Ryan Othniel Kearns

Methodological research rarely generates a broad interest, yet our work on the validity of cluster inference methods for functional magnetic resonance imaging (fMRI) created intense discussion on both the minutia of our approach and its…

Applications · Statistics 2019-01-03 Anders Eklund , Hans Knutsson , Thomas E Nichols

Ensemble clustering aggregates multiple weak clusterings to achieve a more accurate and robust consensus result. The Co-Association matrix (CA matrix) based method is the mainstream ensemble clustering approach that constructs the…

Machine Learning · Computer Science 2024-11-05 Xu Zhang , Yuheng Jia , Mofei Song , Ran Wang

Hierarchical clustering is a stronger extension of one of today's most influential unsupervised learning methods: clustering. The goal of this method is to create a hierarchy of clusters, thus constructing cluster evolutionary history and…

Data Structures and Algorithms · Computer Science 2021-01-14 MohammadTaghi Hajiaghayi , Marina Knittel

Clustering is an unsupervised technique of Data Mining. It means grouping similar objects together and separating the dissimilar ones. Each object in the data set is assigned a class label in the clustering process using a distance measure.…

Information Retrieval · Computer Science 2011-10-13 Parul Agarwal , M. Afshar Alam , Ranjit Biswas

Accuracy, variationality, and convergence underpin the reliability of modern electronic structure methods, yet definitive benchmarks in the relativistic regime remain elusive due to the absence of numerically exact full configuration…

Chemical Physics · Physics 2026-04-03 Shiv Upadhyay , Agam Shayit , Tianyuan Zhang , Stephen H. Yuwono , A. Eugene DePrince , Xiaosong Li

We show how clustering standard errors in one or more dimensions can be justified in M-estimation when there is sampling or assignment uncertainty. Since existing procedures for variance estimation are either conservative or invalid, we…

Econometrics · Economics 2024-11-21 Ruonan Xu , Luther Yap

Spectral Clustering(SC) is a prominent data clustering technique of recent times which has attracted much attention from researchers. It is a highly data-driven method and makes no strict assumptions on the structure of the data to be…

Machine Learning · Computer Science 2019-09-18 Lalith Srikanth Chintalapati , Raghunatha Sarma Rachakonda

It has been noticed that some external CVIs exhibit a preferential bias towards a larger or smaller number of clusters which is monotonic (directly or inversely) in the number of clusters in candidate partitions. This type of bias is caused…

Machine Learning · Statistics 2016-06-21 Yang Lei , James C. Bezdek , Simone Romano , Nguyen Xuan Vinh , Jeffrey Chan , James Bailey

Clustering is spotting pattern in a group of objects and resultantly grouping the similar objects together. Objects have attributes which are not always numerical, sometimes attributes have domain or categories to which they could belong…

Machine Learning · Computer Science 2020-11-20 Utkarsh Nath , Shikha Asrani , Rahul Katarya

Determining the best partition for a dataset can be a challenging task because of 1) the lack of a priori information within an unsupervised learning framework; and 2) the absence of a unique clustering validation approach to evaluate…

Machine Learning · Computer Science 2021-04-06 Isotta Landi , Veronica Mandelli , Michael V. Lombardo

We address the problem of validating the ouput of clustering algorithms. Given data $\mathcal{D}$ and a partition $\mathcal{C}$ of these data into $K$ clusters, when can we say that the clusters obtained are correct or meaningful for the…

Machine Learning · Statistics 2023-02-02 Marina Meilă , Hanyu Zhang

Most existing causal structure learning methods assume data collected from one environment and independent and identically distributed (i.i.d.). In some cases, data are collected from different subjects from multiple environments, which…

Machine Learning · Computer Science 2023-02-07 Wei Chen , Yunjin Wu , Ruichu Cai , Yueguo Chen , Zhifeng Hao

As experimental methods become increasingly popular in crowd dynamics, the question of experimental validity becomes more critical than ever. For the purposes of evaluation of crowd experiments, it is paramount that we distinguish between…

Physics and Society · Physics 2021-12-21 Milad Haghani

Subspace clustering refers to the problem of clustering high-dimensional data into a union of low-dimensional subspaces. Current subspace clustering approaches are usually based on a two-stage framework. In the first stage, an affinity…

Machine Learning · Computer Science 2019-10-22 Shuai Yang , Wenqi Zhu , Yuesheng Zhu

This paper explores the use of clustering methods and machine learning algorithms, including Natural Language Processing (NLP), to identify and classify problems identified in credit risk models through textual information contained in…

Machine Learning · Computer Science 2023-06-05 Szymon Lis , Mariusz Kubkowski , Olimpia Borkowska , Dobromił Serwa , Jarosław Kurpanik