Related papers: Characteristic Length and Clustering
Clustering is a fundamental task in unsupervised learning. Previous research has focused on learning-augmented $k$-means in Euclidean metrics, limiting its applicability to complex data representations. In this paper, we generalize…
A/B testing is a standard approach for evaluating the effect of online experiments; the goal is to estimate the `average treatment effect' of a new feature or condition by exposing a sample of the overall population to it. A drawback with…
Tools of Topological Data Analysis provide stable summaries encapsulating the shape of the considered data. Persistent homology, the most standard and well studied data summary, suffers a number of limitations; its computations are hard to…
In a standard cluster analysis, such as k-means, in addition to clusters locations and distances between them, it's important to know if they are connected or well separated from each other. The main focus of this paper is discovering the…
For a given graph $H$, its subdivisions carry the same topological structure. The existence of $H$-subdivisions within a graph $G$ has deep connections with topological, structural and extremal properties of $G$. One prominent example of…
Motivated by examples from extreme value theory we introduce the general notion of a cluster process as a limiting point process of returns of a certain event in a time series. We explore general invariance properties of cluster processes…
Similarity notions between vertices in a graph, such as structural and regular equivalence, are one of the main ingredients in clustering tools in complex network science. We generalise structural and regular equivalences for undirected…
We propose a new model of cluster growth according to which the probability that a new unit is placed in a point at a distance $r$ from the city center is a Gaussian with mean equal to the cluster radius and variance proportional to the…
In the second article of this series, we establish the convergence of the loop ensemble of interfaces in the random cluster Ising model to a conformal loop ensemble (CLE) --- thus completely describing the scaling limit of the model in…
It has been shown that many complex networks shared distinctive features, which differ in many ways from the random and the regular networks. Although these features capture important characteristics of complex networks, their applicability…
We present new results for LambdaCC and MotifCC, two recently introduced variants of the well-studied correlation clustering problem. Both variants are motivated by applications to network analysis and community detection, and have…
For a harmonic map $u:M^3\to S^1$ on a closed, oriented $3$--manifold, we establish the identity $$2\pi \int_{\theta\in S^1}\chi(\Sigma_{\theta})\geq \frac{1}{2}\int_{\theta\in S^1}\int_{\Sigma_{\theta}}(|du|^{-2}|Hess(u)|^2+R_M)$$ relating…
We develop a simple and unified approach to investigate several aspects of the cluster statistics of random expansive (multi-)sets. In particular, we determine the limiting distribution of the size of the smallest and largest clusters, we…
Understanding the formation and evolution of young star clusters requires quantitative statistical measures of their structure. We investigate the structures of observed and modelled star-forming clusters. By considering the different…
We consider point clouds obtained as random samples of a measure on a Euclidean domain. A graph representing the point cloud is obtained by assigning weights to edges based on the distance between the points they connect. Our goal is to…
We show that if a sample of galaxy clusters is complete above some mass threshold, then hierarchical clustering theories for structure formation predict its autocorrelation function to be determined purely by the cluster abundance and by…
Let $G$ be a finite simple graph. The line graph $L(G)$ represents the adjacencies between edges of $G$. We define first the line simplicial complex $\Delta_L(G)$ of $G$ containing Gallai and anti-Gallai simplicial complexes…
In this paper, we consider feature screening for ultrahigh dimensional clustering analyses. Based on the observation that the marginal distribution of any given feature is a mixture of its conditional distributions in different clusters, we…
Large datasets with interactions between objects are common to numerous scientific fields (i.e. social science, internet, biology...). The interactions naturally define a graph and a common way to explore or summarize such dataset is graph…
We consider random topologies of surfaces generated by cubic interactions. Such surfaces arise in various contexts in 2-dimensional quantum gravity and as world-sheets in string theory. Our results are most conveniently expressed in terms…