Related papers: Cluster Statistics in Expansive Combinatorial Stru…
Clustered data is ubiquitous in a variety of scientific fields. In this paper, we propose a flexible and interpretable modeling approach, called grouped heterogenous mixture modeling, for clustered data, which models cluster-wise…
This paper establishes a combinatorial central limit theorem for stratified randomization, which holds under a Lindeberg-type condition. The theorem allows for an arbitrary number or sizes of strata, with the sole requirement being that…
In presence of long range dispersal, epidemics spread in spatially disconnected regions known as clusters. Here, we characterize exactly their statistical properties in a solvable model, in both the supercritical (outbreak) and critical…
With rapidly increasing data, clustering algorithms are important tools for data analytics in modern research. They have been successfully applied to a wide range of domains; for instance, bioinformatics, speech recognition, and financial…
We formulate a general setting for the cluster expansion method and we discuss sufficient criteria for its convergence. We apply the results to systems of classical and quantum particles with stable interactions.
Determining the number of clusters in a dataset is a fundamental issue in data clustering. Many methods have been proposed to solve the problem of selecting the number of clusters, considering it to be a problem with regard to model…
One key use of k-means clustering is to identify cluster prototypes which can serve as representative points for a dataset. However, a drawback of using k-means cluster centers as representative points is that such points distort the…
Consider a population of $N$ individuals, each having $d\geq 1$ different traits, and an additive measure, called dispersion, which rewards large pairwise separations between traits. The goal is to select $M\leq N$ individuals such that…
We consider a simple model of a growing cluster of points in $\Re^d,d\geq 2$. Beginning with a point $X_1$ located at the origin, we generate a random sequence of points $X_1,X_2,\ldots,X_i,\ldots,$. To generate $X_{i},i\geq 2$ we choose a…
This paper investigates two fundamental descriptors of data, i.e., density distribution versus mass distribution, in the context of clustering. Density distribution has been the de facto descriptor of data distribution since the…
Clustering is one of the most fundamental problems in data analysis and it has been studied extensively in the literature. Though many clustering algorithms have been proposed, clustering theories that justify the use of these clustering…
We develop a novel clustering method for distributional data, where each data point is regarded as a probability distribution on the real line. For distributional data, it has been challenging to develop a clustering method that utilizes…
The present report extends the method of fixed point clustering (Phys.Rev. E 61,5, R4691-4693, 2000) by introducing an indirect criterion for the number of clusters. The derived probability function allows an objective distinction of…
Policy-makers are often faced with the task of distributing a limited supply of resources. To support decision-making in these settings, statisticians are confronted with two challenges: estimands are defined by allocation strategies that…
We present exact and asymptotic results for clusters in the one-dimensional totally asymmetric exclusion process (TASEP) with two different dynamics. The expected length of the largest cluster is shown to diverge logarithmically with…
Sum-of-norms clustering is a convex optimization problem whose solution can be used for the clustering of multivariate data. We propose and study a localized version of this method, and show in particular that it can separate arbitrarily…
The determination of cluster centers generally depends on the scale that we use to analyze the data to be clustered. Inappropriate scale usually leads to unreasonable cluster centers and thus unreasonable results. In this study, we first…
Clustering is one of the most common unsupervised learning tasks in machine learning and data mining. Clustering algorithms have been used in a plethora of applications across several scientific fields. However, there has been limited…
The problem of dimension reduction is of increasing importance in modern data analysis. In this paper, we consider modeling the collection of points in a high dimensional space as a union of low dimensional subspaces. In particular we…
We present a new algorithm for clustering longitudinal data. Data of this type can be conceptualized as consisting of individuals and, for each such individual, observations of a time-dependent variable made at various times. Generically,…