Related papers: Inside the clustering window for random linear equ…
Using an ensemble of high resolution 2D numerical simulations, we explore the scaling properties of cosmological density fluctuations in the non-linear regime. We study the scaling behaviour of the usual $N$--point volume-averaged…
Given a CNF formula F on n variables, the problem of model counting or #SAT is to compute the number of satisfying assignments of F . Model counting is a fundamental but hard problem in computer science with varied applications. Recent…
We analyze the problem of a quantum computer in a correlated environment protected from decoherence by QEC using a perturbative renormalization group approach. The scaling equation obtained reflects the competition between the dimension of…
We propose a Fourier-based approach for optimization of several clustering algorithms. Mathematically, clusters data can be described by a density function represented by the Dirac mixture distribution. The density function can be smoothed…
We study the problem of constructing coresets for clustering problems with time series data. This problem has gained importance across many fields including biology, medicine, and economics due to the proliferation of sensors facilitating…
The problem of CSP sparsification asks: for a given CSP instance, what is the sparsest possible reweighting such that for every possible assignment to the instance, the number of satisfied constraints is preserved up to a factor of $1 \pm…
The standard central limit theorem with a Gaussian attractor for the sum of independent random variables may lose its validity in presence of strong correlations between the added random contributions. Here, we study this problem for…
The generalization gap of a classifier is related to the complexity of the set of functions among which the classifier is chosen. We study a family of low-complexity classifiers consisting of thresholding a random one-dimensional feature.…
Consider a $N\times n$ random matrix $Y_n=(Y_{ij}^{n})$ where the entries are given by $$ Y_{ij}^{n}=\frac{\sigma_{ij}(n)}{\sqrt{n}} X_{ij}^{n} $$ the $X_{ij}^{n}$ being centered, independent and identically distributed random variables…
In this paper we present a generalized configuration model with random triadic closure (GCTC). This model possesses five fundamental properties: large clustering coefficient, power law degree distribution, short path length, non-zero…
Dense granular clusters often behave like macro-particles. We address this interesting phenomenon in a model system of inelastically colliding hard disks inside a circular box, driven by a thermal wall at zero gravity. Molecular dynamics…
We continue the investigation of problems concerning correlation clustering or clustering with qualitative information, which is a clustering formulation that has been studied recently. The basic setup here is that we are given as input a…
The overwhelming majority of empirical research that uses cluster-robust inference assumes that the clustering structure is known, even though there are often several possible ways in which a dataset could be clustered. We propose two tests…
We initiate the study of coresets for clustering in graph metrics, i.e., the shortest-path metric of edge-weighted graphs. Such clustering problems are essential to data analysis and used for example in road networks and data visualization.…
The Central Limit Theorem (CLT) is one of the most fundamental results in statistics. It states that the standardized sample mean of a sequence of $n$ mutually independent and identically distributed random variables with finite first and…
Continuous-variable Gaussian cluster states are a potential resource for universal quantum computation. They can be efficiently and unconditionally built from sources of squeezed light using beam splitters. Here we report on the generation…
We consider the problem of Gaussian mixture clustering in the high-dimensional limit where the data consists of $m$ points in $n$ dimensions, $n,m \rightarrow \infty$ and $\alpha = m/n$ stays finite. Using exact but non-rigorous methods…
We consider estimation in a high-dimensional linear model with strongly correlated variables. We propose to cluster the variables first and do subsequent sparse estimation such as the Lasso for cluster-representatives or the group Lasso…
The basic random $k$-SAT problem is: Given a set of $n$ Boolean variables, and $m$ clauses of size $k$ picked uniformly at random from the set of all such clauses on our variables, is the conjunction of these clauses satisfiable? Here we…
In clustering problems, a central decision-maker is given a complete metric graph over vertices and must provide a clustering of vertices that minimizes some objective function. In fair clustering problems, vertices are endowed with a color…