Related papers: Non-Reconstructability in the Stochastic Block Mod…
We show that the standard 1+1d $\mathbb{Z}_2\times \mathbb{Z}_2$ cluster model has a non-invertible global symmetry, described by the fusion category Rep(D$_8$). Therefore, the cluster state is not only a $\mathbb{Z}_2\times \mathbb{Z}_2$…
We study the cluster recovery problem in the semi-supervised active clustering framework. Given a finite set of input points, and an oracle revealing whether any two points lie in the same cluster, our goal is to recover all clusters…
We study statistical and computational limits of clustering when the means of the centres are sparse and their dimension is possibly much larger than the sample size. Our theoretical analysis focuses on the model $X_i = z_i \theta +…
When there is no independence, abnormal observations may have a tendency to appear in clusters instead of scattered along the time frame. Identifying clusters and estimating their size are important problems arising in statistics of…
Many clustering algorithms are guided by certain cost functions such as the widely-used $k$-means cost. These algorithms divide data points into clusters with often complicated boundaries, creating difficulties in explaining the clustering…
Classical blockmodel is known as the simplest among models of networks with community structure. The model can be also seen as an extremely simply example of interconnected networks. For this reason, it is surprising that the percolation…
We consider a simple model of a growing cluster of points in $\Re^d,d\geq 2$. Beginning with a point $X_1$ located at the origin, we generate a random sequence of points $X_1,X_2,\ldots,X_i,\ldots,$. To generate $X_{i},i\geq 2$ we choose a…
Stochastic block models (SBMs) are a very commonly studied network model for community detection algorithms. In the standard form of an SBM, the $n$ vertices (or nodes) of a graph are generally divided into multiple pre-determined…
In empirical work it is common to estimate parameters of models and report associated standard errors that account for "clustering" of units, where clusters are defined by factors such as geography. Clustering adjustments are typically…
We study supercritical branching processes in which all particles evolve according to some general Markovian motion (which may possess absorbing states) and branch independently at a fixed constant rate. Under fairly natural assumptions on…
Non-equilibrium cluster-cluster aggregation of particles diffusing in or at the cell membrane has been hypothesized to lead to domains of finite size in different biological contexts such as lipid rafts, cell adhesion complexes, or…
We consider a bipartite stochastic block model on vertex sets $V_1$ and $V_2$, with planted partitions in each, and ask at what densities efficient algorithms can recover the partition of the smaller vertex set. When $|V_2| \gg |V_1|$,…
This paper investigates graph clustering in the planted cluster model in the presence of {\em small clusters}. Traditional results dictate that for an algorithm to provably correctly recover the clusters, {\em all} clusters must be…
The problem of reconstructing a sequence of independent and identically distributed symbols from a set of equal size, consecutive, fragments, as well as a dependent reference sequence, is considered. First, in the regime in which the…
We present a general approach to the bulk-boundary correspondence of noninvertible topological phases, including both topological and fracton orders. This is achieved by a novel bulk construction protocol where solvable $(d+1)$-dimensional…
We consider branching processes for structured populations: each individual is characterized by a type or trait which belongs to a general measurable state space. We focus on the supercritical recurrent case, where the population may…
Networks, which represent agents and interactions between them, arise in myriad applications throughout the sciences, engineering, and even the humanities. To understand large-scale structure in a network, a common task is to cluster a…
Block coordinate descent methods and stochastic subgradient methods have been extensively studied in optimization and machine learning. By combining randomized block sampling with stochastic subgradient methods based on dual averaging, we…
We use a measure of clustering derived from the nearest neighbour distribution and the void probability function to distinguish between regular and clustered structures. This measure offers a succinct way to incorporate additional…
In mixture modeling and clustering applications, the number of components and clusters is often not known. A stick-breaking mixture model, such as the Dirichlet process mixture model, is an appealing construction that assumes infinitely…