English
Related papers

Related papers: A comparison of Gap statistic definitions with and…

200 papers

Generalized $k$-means can be incorporated with any similarity or dissimilarity measure for clustering. By choosing the dissimilarity measure as the well known likelihood ratio or $F$-statistic, this work proposes a method based on…

Methodology · Statistics 2020-08-11 Tonglin Zhang , Ge Lin

A computational theory for clustering and a semi-supervised clustering algorithm is presented. Clustering is defined to be the obtainment of groupings of data such that each group contains no anomalies with respect to a chosen grouping…

Machine Learning · Computer Science 2025-07-17 Nassir Mohammad

We present exact and asymptotic results for clusters in the one-dimensional totally asymmetric exclusion process (TASEP) with two different dynamics. The expected length of the largest cluster is shown to diverge logarithmically with…

Statistical Mechanics · Physics 2009-11-07 O. Pulkkinen , J. Merikoski

Cluster validity indexes are very important tools designed for two purposes: comparing the performance of clustering algorithms and determining the number of clusters that best fits the data. These indexes are in general constructed by…

Machine Learning · Computer Science 2018-12-24 Ahmed Ben Said , Rachid Hadjidj , Sebti Foufou

Probability density functions (PDF) of statistical distributions of cluster sizes N, where N is the number of particles in the cluster, often seem to have less freedom than expected from considering the number of degrees of freedom at the…

Data Analysis, Statistics and Probability · Physics 2011-03-08 Sascha Vongehr , Shaochun Tang , Xiangkang Meng

The objective of clustering is to discover natural groups in datasets and to identify geometrical structures which might reside there, without assuming any prior knowledge on the characteristics of the data. The problem can be seen as…

Computational Geometry · Computer Science 2018-01-26 Luis-Evaristo Caraballo , José-Miguel Díaz-Báñez , Nadine Kroher

The problem of finding groups in data (cluster analysis) has been extensively studied by researchers from the fields of Statistics and Computer Science, among others. However, despite its popularity it is widely recognized that the…

Statistics Theory · Mathematics 2013-10-10 José E. Chacón

Comparing the representations learned by different neural networks has recently emerged as a key tool to understand various architectures and ultimately optimize them. In this work, we introduce GULP, a family of distance measures between…

Machine Learning · Computer Science 2022-10-14 Enric Boix-Adsera , Hannah Lawrence , George Stepaniants , Philippe Rigollet

Consider $n$ independent measurements, with the additional information of the times at which measurements are performed. This paper deals with testing statistical hypotheses when $n$ is large and only a small amount of observations…

Methodology · Statistics 2018-09-26 Emanuele Dolera , Stefano Favaro , Andrea Bulgarelli , Alessio Aboudan

The $k$-means algorithm is arguably the most popular nonparametric clustering method but cannot generally be applied to datasets with incomplete records. The usual practice then is to either impute missing values under an assumed…

Machine Learning · Statistics 2018-09-11 Andrew Lithio , Ranjan Maitra

We consider the probability of having two intervals (gaps) without eigenvalues in the bulk scaling limit of the Gaussian Unitary Ensemble of random matrices. We describe uniform asymptotics for the transition between a single large gap and…

Functional Analysis · Mathematics 2020-03-19 Benjamin Fahs , Igor Krasovsky

Clustering is one of the most fundamental and wide-spread techniques in exploratory data analysis. Yet, the basic approach to clustering has not really changed: a practitioner hand-picks a task-specific clustering loss to optimize and fit…

Machine Learning · Computer Science 2019-11-01 Yibo Jiang , Nakul Verma

We consider semi-supervised classification when part of the available data is unlabeled. These unlabeled data can be useful for the classification problem when we make an assumption relating the behavior of the regression function to that…

Statistics Theory · Mathematics 2007-06-13 Philippe Rigollet

In these notes we describe heuristics to predict computational-to-statistical gaps in certain statistical problems. These are regimes in which the underlying statistical problem is information-theoretically possible although no efficient…

Machine Learning · Statistics 2018-04-23 Afonso S. Bandeira , Amelia Perry , Alexander S. Wein

The global clustering coefficient is an effective measure for analyzing and comparing the structures of complex networks. The random annulus graph is a modified version of the well-known Erd\H{o}s-R\'{e}nyi random graph. It has been…

Methodology · Statistics 2025-10-20 Mingao Yuan

Teramoto et al. defined a new measure called the gap ratio that measures the uniformity of a finite point set sampled from $\cal S$, a bounded subset of $\mathbb{R}^2$. We generalize this definition of measure over all metric spaces by…

Computational Geometry · Computer Science 2015-12-07 Arijit Bishnu , Sameer Desai , Arijit Ghosh , Mayank Goswami , Subhabrata Paul

There is no, nor will there ever be, single best clustering algorithm. Nevertheless, we would still like to be able to distinguish between methods that work well on certain task types and those that systematically underperform. Clustering…

Machine Learning · Computer Science 2025-10-16 Marek Gagolewski

Let a cluster be a term with a number of patterns occurring in it. We give two accounts of clusters, a geometric one as sets of (node and edge) positions, and an inductive one as pairs of terms with gaps (2nd order variables) and…

Logic in Computer Science · Computer Science 2017-08-29 Nao Hirokawa , Julian Nagele , Vincent van Oostrom , Michio Oyamaguchi

The distribution function of particles over clusters is proposed for a system of identical intersecting spheres, the centres of which are uniformly distributed in space. Consideration is based on the concept of the rank number of clusters,…

Statistical Mechanics · Physics 2023-02-01 Murat Kh. Khokonov , Azamat Kh. Khokonov

The logistic map is a nonlinear difference equation well studied in the literature, used to model self-limiting growth in certain populations. It is known that, under certain regularity conditions, the stochastic logistic map, where the…

Dynamical Systems · Mathematics 2023-10-05 Maricela Cruz , Austin Wei , Johanna Hardin , Ami Radunskaya