Related papers: Risk Bounds For Mode Clustering
Bayesian clustering typically relies on mixture models, with each component interpreted as a different cluster. After defining a prior for the component parameters and weights, Markov chain Monte Carlo (MCMC) algorithms are commonly used to…
Quality assessments of models in unsupervised learning and clustering verification in particular have been a long-standing problem in the machine learning research. The lack of robust and universally applicable cluster validity scores often…
Detection of change-points in a sequence of high-dimensional observations is a very challenging problem, and this becomes even more challenging when the sample size (i.e., the sequence length) is small. In this article, we propose some…
We review recent advancements in cosmology with galaxy clusters. Galaxy clusters are the most massive objects in the Universe. Consequently the cluster number density as a function of cluster mass, or cluster abundance, is sensitive to…
The scaling properties of the cluster size distribution of a system of diffusing clusters is studied in terms of a simple kinetic mean field model. It is shown that a one parameter family of mathematically valid scaling solutions exists.…
Semi-supervised clustering methods incorporate a limited amount of supervision into the clustering process. Typically, this supervision is provided by the user in the form of pairwise constraints. Existing methods use such constraints in…
An important issue in clustering concerns the avoidance of false positives while searching for clusters. This work addressed this problem considering agglomerative methods, namely single, average, median, complete, centroid and Ward's…
In this article, we propose the use of partitioning and clustering methods as an alternative to Gaussian quadrature for stochastic collocation. The key idea is to use cluster centers as the nodes for collocation. In this way, we can extend…
We study the relative fraction of galaxy morphological types in clusters, as a function of the projected local galaxy density and different global parameters: cluster projected gas density, cluster projected total mass density , and reduced…
Cluster analysis is an unsupervised learning strategy that can be employed to identify subgroups of observations in data sets of unknown structure. This strategy is particularly useful for analyzing high-dimensional data such as microarray…
In this paper, we test whether two datasets share a common clustering structure. As a leading example, we focus on comparing clustering structures in two independent random samples from two mixtures of multivariate normal distributions.…
In the modal approach to clustering, clusters are defined as the local maxima of the underlying probability density function, where the latter can be estimated either non-parametrically or using finite mixture models. Thus, clusters are…
Many high dimensional vector distances tend to a constant. This is typically considered a negative "contrast-loss" phenomenon that hinders clustering and other machine learning techniques. We reinterpret "contrast-loss" as a blessing.…
Cluster growth in a coagulating system of active particles (such as microswimmers in a solvent) is studied by theory and simulation. In contrast to passive systems, the net velocity of a cluster can have various scalings dependent on the…
We propose a scheme for parameter estimation with cluster states. We find that phase estimation with cluster states under a many-body Hamiltonian and separable measurements leads to a precision at the Heisenberg limit. As noise models we…
The increasing needs of clustering massive datasets and the high cost of running clustering algorithms poses difficult problems for users. In this context it is important to determine if a data set is clusterable, that is, it may be…
In empirical work it is common to estimate parameters of models and report associated standard errors that account for "clustering" of units, where clusters are defined by factors such as geography. Clustering adjustments are typically…
Scatterplots are used for a variety of visual analytics tasks, including cluster identification, and the visual encodings used on a scatterplot play a deciding role on the level of visual separation of clusters. For visualization designers,…
The near-threshold clustering phenomenon is well understood by the low-energy universality, for shallow bound states below the threshold. Nevertheless, the characteristics of resonances slightly above the threshold still lack thorough…
The clustering of a data set is one of the core tasks in data analytics. Many clustering algorithms exhibit a strong contrast between a favorable performance in practice and bad theoretical worst-cases. Prime examples are least-squares…