Related papers: Stability of Density-Based Clustering
Predictions are often probabilities; e.g., a prediction could be for precipitation tomorrow, but with only a 30% chance. Given such probabilistic predictions together with the actual outcomes, "reliability diagrams" help detect and diagnose…
In this paper we propose a measure of clustering quality or accuracy that is appropriate in situations where it is desirable to evaluate a clustering algorithm by somehow comparing the clusters it produces with ``ground truth' consisting of…
Gravitational clustering in the nonlinear regime remains poorly understood. Gravity dual of gravitational clustering has recently been proposed as a means to study the nonlinear regime. The stable clustering ansatz remains a key ingredient…
We study the problem of estimating the density $f(\boldsymbol x)$ of a random vector ${\boldsymbol X}$ in $\mathbb R^d$. For a spanning tree $T$ defined on the vertex set $\{1,\dots ,d\}$, the tree density $f_{T}$ is a product of bivariate…
Traditionally, the Dirichlet-multinomial distribution has been recognized as a key model for contingency tables generated by cluster sampling schemes. There are, however, other possible distributions appropriate for these contingency…
We address the question of how well the density profile of galaxy clusters can be determined by combining strong lensing and velocity dispersion data. We use cosmological dark matter simulations of clusters to test the reliability of the…
We use the presently observed number density of large X-ray clusters and linear mass power spectra to constrain the shape parameter ($\Gamma$), the spectral index ($n$), the amplitude of matter density perturbations on the scale of $8…
Clustering is a data analysis method for extracting knowledge by discovering groups of data called clusters. Among these methods, state-of-the-art density-based clustering methods have proven to be effective for arbitrary-shaped clusters.…
When some 'entities' are related by the 'features' they share they are amenable to a bipartite network representation. Plant-pollinator ecological communities, co-authorship of scientific papers, customers and purchases, or answers in a…
Measurements of the cluster abundance as a function of mass and redshift provide an important cosmological test that probe not only the expansion rate but also the growth of perturbations. In this paper we adopt a scalar field scenario…
We consider the relative configurational entropy per cell S_Delta as a measure of the degree of spatial disorder for systems of finite-sized objects. It is highly sensitive to deviations from the most spatially ordered reference…
The two most extended density-based approaches to clustering are surely mixture model clustering and modal clustering. In the mixture model approach, the density is represented as a mixture and clusters are associated to the different…
Density-based clustering methodology has been widely considered in the statistical literature for classifying Euclidean observations. However, this approach has not been contemplated for directional data yet. In this work, directional…
Latent variable models for network data extract a summary of the relational structure underlying an observed network. The simplest possible models subdivide nodes of the network into clusters; the probability of a link between any two nodes…
Density mode clustering is a nonparametric clustering method. The clusters are the basins of attraction of the modes of a density estimator. We study the risk of mode-based clustering. We show that the clustering risk over the cluster cores…
We investigate a critical scaling law for the cluster heterogeneity $H$ in site and bond percolations in $d$-dimensional lattices with $d=2,...,6$. The cluster heterogeneity is defined as the number of distinct cluster sizes. As an…
A statistical description of heavy particles suspended in incompressible rough self-similar flows is developed. It is shown that, differently from smooth flows, particles do not form fractal clusters. They rather distribute inhomogeneously…
Clustering is an essential technique for discovering patterns in data. The steady increase in amount and complexity of data over the years led to improvements and development of new clustering algorithms. However, algorithms that can…
This paper presents and analyzes an approach to cluster-based inference for dependent data. The primary setting considered here is with spatially indexed data in which the dependence structure of observed random variables is characterized…
We present a nonparametric method for selecting informative features in high-dimensional clustering problems. We start with a screening step that uses a test for multimodality. Then we apply kernel density estimation and mode clustering to…