Related papers: Regularity in the distribution of superclusters?
An automatic procedure to perform sub-clustering on large samples is presented. At each iteration, the most diverse cluster is sub-clustered, and the global diversity of the new classification is compared to the previous one. The process…
Nonparametric Bayesian approaches provide a flexible framework for clustering without pre-specifying the number of groups, yet they are well known to overestimate the number of clusters, especially for functional data. We show that a…
Spectral clustering requires the time-consuming decomposition of the Laplacian matrix of the similarity graph, thus limiting its applicability to large datasets. To improve the efficiency of spectral clustering, a top-down approach was…
Data clustering is the process of identifying natural groupings or clusters within multidimensional data based on some similarity measure. Clustering is a fundamental process in many different disciplines. Hence, researchers from different…
Substructure in galaxy clusters can be quantified with the robust Delta statistics (Dressler and Shectman 1988) which uses velocity kinematics and sky projected positions. We test its sensitivity using dissipationless numerical simulations…
Based on an expert systems approach, the issue of community detection can be conceptualized as a clustering model for networks. Building upon this further, community structure can be measured through a clustering coefficient, which is…
We analyse the dependence of clustering properties of galaxies as a function of their large-scale environment. In order to characterize the environment on large scales, we use the catalogue of future virialized superstructures (FVS) by…
Clustering in image analysis is a central technique that allows to classify elements of an image. We describe a simple clustering technique that uses the method of similarity matrices. We expand upon recent results in spectral analysis for…
Biclustering involves the simultaneous clustering of objects and their attributes, thus defining local two-way clustering models. Recently, efficient algorithms were conceived to enumerate all biclusters in real-valued datasets. In this…
We develop a network in which the natural numbers are the vertices. We use the decomposition of natural numbers by prime numbers to establish the connections. We perform data collapse and show that the degree distribution of these networks…
The analysis of the presence of substructures in 16 well-sampled clusters of galaxies suggests a stimulating hypothesis: Clusters could be classified as unimodal or bimodal, on the basis of to the sub-clump distribution in the {\em 3-D}…
Stars form predominantly in groups usually denoted as clusters or associations. The observed stellar groups display a broad spectrum of masses, sizes and other properties, so it is often assumed that there is no underlying structure in this…
Clusters of galaxies are often embedded in larger-scale superclusters with dimensions of tens or perhaps even hundreds of Mpc. Observational and theoretical evidence suggest an important connection between cluster properties and their…
Traditionally, the Dirichlet-multinomial distribution has been recognized as a key model for contingency tables generated by cluster sampling schemes. There are, however, other possible distributions appropriate for these contingency…
An unsupervised classification method for point events occurring on a network of lines is proposed. The idea relies on the distributional flexibility and practicality of random partition models to discover the clustering structure featuring…
In an age of increasingly large data sets, investigators in many different disciplines have turned to clustering as a tool for data analysis and exploration. Existing clustering methods, however, typically depend on several nontrivial…
While clustering is ubiquitously used across science and industry, uncertainty in cluster assignments is rarely quantified with rigorous guarantees. We propose a novel conformal inference framework for clustering that returns confidence…
Satellites within simulated massive clusters are significantly spatially correlated with each other, even when those satellites are not gravitationally bound to each other. This correlation is produced by satellites that entered their hosts…
The statistical behavior of the size (or mass) of the largest cluster in subcritical percolation on a finite lattice of size $N$ is investigated (below the upper critical dimension, presumably $d_c=6$). It is argued that as $N \to \infty$…
The consistency of the maximum likelihood estimator for mixtures of elliptically-symmetric distributions for estimating its population version is shown, where the underlying distribution $P$ is nonparametric and does not necessarily belong…