Related papers: Relative cluster entropy for power-law correlated …
Correlation clustering is a central problem in unsupervised learning, with applications spanning community detection, duplicate detection, automated labelling and many more. In the correlation clustering problem one receives as input a set…
Clustering is a widely used unsupervised learning method for finding structure in the data. However, the resulting clusters are typically presented without any guarantees on their robustness; slightly changing the used data sample or…
We show, on purely statistical grounds and without appeal to any physical model, that a power-law $q-$entropy $S_q$, with $0<q<1$, can be {\it extensive}. More specifically, if the components $X_i$ of a vector $X \in \mathbb{R}^N$ are…
A major challenge in cluster analysis is that the number of data clusters is mostly unknown and it must be estimated prior to clustering the observed data. In real-world applications, the observed data is often subject to heavy tailed noise…
In simple colloidal suspensions, clusters are various multimers that result from colloid self-association and exist in equilibrium with monomers.There are two types of potentials that are known to produce clusters: a) potentials that result…
Mixture models and topic models generate each observation from a single cluster, but standard variational posteriors for each observation assign positive probability to all possible clusters. This requires dense storage and runtime costs…
Although some information-theoretic measures of uncertainty or granularity have been proposed in rough set theory, these measures are only dependent on the underlying partition and the cardinality of the universe, independent of the lower…
This paper presents and analyzes an approach to cluster-based inference for dependent data. The primary setting considered here is with spatially indexed data in which the dependence structure of observed random variables is characterized…
This study investigates empirically whether the degree of stock market efficiency is related to the prediction power of future price change using the indices of twenty seven stock markets. Efficiency refers to weak-form efficient market…
Finding a good clustering of vertices in a network, where vertices in the same cluster are more tightly connected than those in different clusters, is a useful, important, and well-studied task. Many clustering algorithms scale well,…
Clustering analysis identifies samples as groups based on either their mutual closeness or homogeneity. In order to detect clusters in arbitrary shapes, a novel and generic solution based on boundary erosion is proposed. The clusters are…
The clustering of small heavy inertial particles subjected to Stokes drag in turbulence is known to be minimal at small and large Stokes number and substantial at $\rm St = \mathcal O(1)$. This non-monotonic trend, which has been shown…
This work initiates the study of memory-query tradeoffs for graph problems, with a focus on correlation clustering. Correlation clustering asks for a partition of the vertices that minimizes disagreements: non-edges inside clusters plus…
The recent quantum information boom has effected a resurgence of interest in unitary coupled cluster (UCC) theory. Our group's interest in local energy landscapes of unitary ans\"atze prompted us to investigate the classical approach of…
Herein, we propose a site random cluster model by introducing an additional cluster weight in the partition function of the traditional site percolation. To simulate the model on a square lattice, we combine the color-assignation and the…
We present an analysis of different sets of gravitational N-body simulations, all describing the dynamics of discrete particles with a small initial velocity dispersion. They encompass very different initial particle configurations,…
Relative entropy is a measure of distinguishability for quantum states, and plays a central role in quantum information theory. The family of Renyi entropies generalizes to Renyi relative entropies that include as special cases most entropy…
The correct identification of clusters is crucial for an accurate monitoring of the spread of a disease and also in many other natural, social and physical phenomena which exhibit an epidemic structure. Nevertheless, even when an accurate…
In this paper, we present a local information theoretic approach to explicitly learn probabilistic clustering of a discrete random variable. Our formulation yields a convex maximization problem for which it is NP-hard to find the global…
We develop a cluster typical medium theory to study localization in disordered electronic systems. Our formalism is able to incorporate non-local correlations beyond the local typical medium theory in a systematic way. The cluster typical…