Related papers: Cheeger $N$-clusters
A $d$-simplex is defined to be a collection $A_1,\dots,A_{d+1}$ of subsets of size $k$ of $[n]$ such that the intersection of all of them is empty, but the intersection of any $d$ of them is non-empty. Furthermore, a $d$-cluster is a…
The Bayesian approach to inference stands out for naturally allowing borrowing information across heterogeneous populations, with different samples possibly sharing the same distribution. A popular Bayesian nonparametric model for…
We study an infinite system of ordinary differential equations that models the evolution of coagulating and fragmenting clusters, which we assume to be composed of identical units. Under very mild assumptions on the coefficients we prove…
An unsupervised classification method for point events occurring on a network of lines is proposed. The idea relies on the distributional flexibility and practicality of random partition models to discover the clustering structure featuring…
In this work, we prove an existence result for an optimal partition problem of the form $$\min \{F_s(A_1,\dots,A_m)\colon A_i \in \mathcal{A}_s, \, A_i\cap A_j =\emptyset \mbox{ for } i\neq j\},$$ where $F_s$ is a cost functional with…
Locally isoperimetric $N$-partitions are partitions of the space $\mathbb R^d$ into $N$ regions with prescribed, finite or infinite measure, which have minimal perimeter (which is the $(d-1)$-dimensional measure of the interfaces between…
Improving the theoretical description of galaxy clustering on small scales is an important challenge in cosmology, as it can considerably increase the scientific return of forthcoming galaxy surveys -- e.g. tightening the bounds on neutrino…
We study density thresholds that force a measurable set $E\subseteq\mathbb{R}^d$ to contain all sufficiently large similar copies of every $n$-point configuration. We prove a lower bound of the form $1-O((\log n)/n)$, which matches the…
We propose a new method for clustering based on the local minimization of the \gamma-divergence, which we call the spontaneous clustering. The greatest advantage of the proposed method is that it automatically detects the number of clusters…
We give necessary and sufficient conditions for a regular semi-Dirichlet form to enjoy a new Feller type property, which we call \emph{weak Feller property}. Our characterization involves potential theoretic as well as probabilistic aspects…
Recently, there has been substantial interest in clustering research that takes a beyond worst-case approach to the analysis of algorithms. The typical idea is to design a clustering algorithm that outputs a near-optimal solution, provided…
Clustering is one of the main tasks in exploratory data analysis and descriptive statistics where the main objective is partitioning observations in groups. Clustering has a broad range of application in varied domains like climate,…
Evidential clustering is an approach to clustering based on the use of Dempster-Shafer mass functions to represent cluster-membership uncertainty. In this paper, we introduce a neural-network based evidential clustering algorithm, called…
Clustering is a fundamental building block of modern statistical analysis pipelines. Fair clustering has seen much attention from the machine learning community in recent years. We are some of the first to study fairness in the context of…
We introduce a novel criterion in clustering that seeks clusters with limited range of values associated with each cluster's elements. In clustering or classification the objective is to partition a set of objects into subsets, called…
In fully-dynamic consistent clustering, we are given a finite metric space $(M,d)$, and a set $F\subseteq M$ of possible locations for opening centers. Data points arrive and depart, and the goal is to maintain an approximately optimal…
Let $G$ be a finitely generated Kleinian group and let $\Delta$ be an invariant collection of components in its region of discontinuity. The Teichm\"uller space $T(\Delta,G)$ supported in $\Delta$, is the space of equivalence classes of…
We define a family of functionals, called p-oscillation functionals, that can be interpreted as discrete versions of the classical total variation functional for p=1 and of the p-Dirichlet functionals for p>1. We introduce the notion of…
We consider the problem of clustering a set of high-dimensional data points into sets of low-dimensional linear subspaces. The number of subspaces, their dimensions, and their orientations are unknown. We propose a simple and low-complexity…
Clustering is a pivotal challenge in unsupervised machine learning and is often investigated through the lens of mixture models. The optimal error rate for recovering cluster labels in Gaussian and sub-Gaussian mixture models involves ad…